{
  "id": 243845,
  "title": "47th Place Solution & Code",
  "url": "/competitions/bms-molecular-translation/discussion/243845",
  "author_name": "Nikita Kozodoi",
  "post_date": "2021-06-04T08:26:48.509000",
  "votes": 31,
  "comment_count": 8,
  "views": 0,
  "content": "<h3>Summary</h3>\n<p>First, let me congratulate all the winners and thank the competition hosts - it was a really interesting challenge! It was great to be a part of a competition requiring both CV and NLP skills and having such a large dataset with a very minimal shakeup. </p>\n<p>The (not so) funny thing is that I was confident that the competition deadline is actually two days later, 05.06. I mistakingly created a calendar entry for the deadline and never checked if it's actually correct. What a surprise for me was to check out the competition page yesterday and realize that it was just 10 hours before the end! 😅</p>\n<p>My solution is an ensemble of seven CNN-LSTM Encoder-Decoder models. All models are implemented in PyTorch and trained on a local machine with Quadro RTX 6000 GPU.</p>\n<h3>Code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/kozodoi/47th-place-solution-bms-ensembling\" target=\"_blank\">Kaggle notebook</a> reproducing my submission and ensembling pipeline</li>\n<li><a href=\"https://github.com/kozodoi/BMS_Molecular_Translation\" target=\"_blank\">GitHub repo</a> with the complete training codes</li>\n</ul>\n<h3>Data</h3>\n<ul>\n<li>single 80/20 train/test split stratified by the molecule length</li>\n<li>creating 3M extra images from <code>extra_approved_InChIs</code> using RDKit (thanks <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a>)</li>\n<li>on each epoch, I was taking all train images + random 1M extra images for training</li>\n</ul>\n<h3>Tokenizer</h3>\n<ul>\n<li>my tokenizer is very similar to the one proposed by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> in his great pipeline</li>\n<li>using training + extra data to fit the tokenizer, so it ended up having about 20 more tokens due to more distinct molecules in the data</li>\n</ul>\n<h3>Image augmentations</h3>\n<ul>\n<li>making lines and letters thicker with cv2 morphological transformations:</li>\n</ul>\n<pre><code>image = cv2.morphologyEx(image, cv2.MORPH_OPEN, np.ones((2, 2)))\nimage = cv2.erode(image, np.ones((2, 2)))\n</code></pre>\n<ul>\n<li>small image rotations during training:</li>\n</ul>\n<pre><code>ShiftScaleRotate(0.01, 0.01, 0.10)\n</code></pre>\n<ul>\n<li>cropping pictures with large empty borders (thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>)</li>\n<li>rotating test images with height &gt; width (thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>)<br>\n<img src=\"https://i.postimg.cc/t4KqLMNC/inchi.jpg\" alt=\"images\"></li>\n</ul>\n<h3>Base models</h3>\n<p>All models used EfficientNet-based CNN encoder and LSTM-based RNN decoder. The table below from Neptune.ai provides the main architecture and training parameters of my top models:<br>\n<img src=\"https://i.postimg.cc/cLrTp1Pc/Screen-2021-06-04-at-10-17-02.jpg\" alt=\"models\"></p>\n<h3>Ensembling &amp; post-processing</h3>\n<ul>\n<li>using beam search with k = 5 on most of the base models (thanks <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a>)</li>\n<li>normalizing each model predictions with RDKit (thanks <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>)</li>\n<li>ensembling predictions from base models using majority voting</li>\n<li>setting \"unconfident\" predictions to the ones produced by the lowest-CV model that has a valid prediction according to RDKit (thanks <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> again)</li>\n</ul>\n<h3>What else I planned to try</h3>\n<ul>\n<li>adding image meta-data such as pixel resolution and height/width ratio as input features</li>\n<li>iteratively predicting different parts of the InChI (formula, /c and /h parts) with separate models that would take image and predictions of the previous molecule part as input</li>\n<li>switching to transformers. I only started working on transformer pipeline a few days ago and did not have enough time to train a heavy model due to limited resources</li>\n</ul>\n<p>The final solution achieves <strong>1.32</strong> on the public and <strong>1.31</strong> on the private LB (47th place). I hope this summary was useful for some of you. Happy to answer any questions in the comments and see you in the next competitions! 😊</p>",
  "messages": [
    {
      "id": 1335454,
      "postDate": "2021-06-04T08:26:48.510Z",
      "content": "<h3>Summary</h3>\n<p>First, let me congratulate all the winners and thank the competition hosts - it was a really interesting challenge! It was great to be a part of a competition requiring both CV and NLP skills and having such a large dataset with a very minimal shakeup. </p>\n<p>The (not so) funny thing is that I was confident that the competition deadline is actually two days later, 05.06. I mistakingly created a calendar entry for the deadline and never checked if it's actually correct. What a surprise for me was to check out the competition page yesterday and realize that it was just 10 hours before the end! 😅</p>\n<p>My solution is an ensemble of seven CNN-LSTM Encoder-Decoder models. All models are implemented in PyTorch and trained on a local machine with Quadro RTX 6000 GPU.</p>\n<h3>Code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/kozodoi/47th-place-solution-bms-ensembling\" target=\"_blank\">Kaggle notebook</a> reproducing my submission and ensembling pipeline</li>\n<li><a href=\"https://github.com/kozodoi/BMS_Molecular_Translation\" target=\"_blank\">GitHub repo</a> with the complete training codes</li>\n</ul>\n<h3>Data</h3>\n<ul>\n<li>single 80/20 train/test split stratified by the molecule length</li>\n<li>creating 3M extra images from <code>extra_approved_InChIs</code> using RDKit (thanks <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a>)</li>\n<li>on each epoch, I was taking all train images + random 1M extra images for training</li>\n</ul>\n<h3>Tokenizer</h3>\n<ul>\n<li>my tokenizer is very similar to the one proposed by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> in his great pipeline</li>\n<li>using training + extra data to fit the tokenizer, so it ended up having about 20 more tokens due to more distinct molecules in the data</li>\n</ul>\n<h3>Image augmentations</h3>\n<ul>\n<li>making lines and letters thicker with cv2 morphological transformations:</li>\n</ul>\n<pre><code>image = cv2.morphologyEx(image, cv2.MORPH_OPEN, np.ones((2, 2)))\nimage = cv2.erode(image, np.ones((2, 2)))\n</code></pre>\n<ul>\n<li>small image rotations during training:</li>\n</ul>\n<pre><code>ShiftScaleRotate(0.01, 0.01, 0.10)\n</code></pre>\n<ul>\n<li>cropping pictures with large empty borders (thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>)</li>\n<li>rotating test images with height &gt; width (thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>)<br>\n<img src=\"https://i.postimg.cc/t4KqLMNC/inchi.jpg\" alt=\"images\"></li>\n</ul>\n<h3>Base models</h3>\n<p>All models used EfficientNet-based CNN encoder and LSTM-based RNN decoder. The table below from Neptune.ai provides the main architecture and training parameters of my top models:<br>\n<img src=\"https://i.postimg.cc/cLrTp1Pc/Screen-2021-06-04-at-10-17-02.jpg\" alt=\"models\"></p>\n<h3>Ensembling &amp; post-processing</h3>\n<ul>\n<li>using beam search with k = 5 on most of the base models (thanks <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a>)</li>\n<li>normalizing each model predictions with RDKit (thanks <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>)</li>\n<li>ensembling predictions from base models using majority voting</li>\n<li>setting \"unconfident\" predictions to the ones produced by the lowest-CV model that has a valid prediction according to RDKit (thanks <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> again)</li>\n</ul>\n<h3>What else I planned to try</h3>\n<ul>\n<li>adding image meta-data such as pixel resolution and height/width ratio as input features</li>\n<li>iteratively predicting different parts of the InChI (formula, /c and /h parts) with separate models that would take image and predictions of the previous molecule part as input</li>\n<li>switching to transformers. I only started working on transformer pipeline a few days ago and did not have enough time to train a heavy model due to limited resources</li>\n</ul>\n<p>The final solution achieves <strong>1.32</strong> on the public and <strong>1.31</strong> on the private LB (47th place). I hope this summary was useful for some of you. Happy to answer any questions in the comments and see you in the next competitions! 😊</p>",
      "rawMarkdown": "### Summary\n\nFirst, let me congratulate all the winners and thank the competition hosts - it was a really interesting challenge! It was great to be a part of a competition requiring both CV and NLP skills and having such a large dataset with a very minimal shakeup. \n\nThe (not so) funny thing is that I was confident that the competition deadline is actually two days later, 05.06. I mistakingly created a calendar entry for the deadline and never checked if it's actually correct. What a surprise for me was to check out the competition page yesterday and realize that it was just 10 hours before the end! 😅\n\nMy solution is an ensemble of seven CNN-LSTM Encoder-Decoder models. All models are implemented in PyTorch and trained on a local machine with Quadro RTX 6000 GPU.\n\n\n### Code\n- [Kaggle notebook](https://www.kaggle.com/kozodoi/47th-place-solution-bms-ensembling) reproducing my submission and ensembling pipeline\n- [GitHub repo](https://github.com/kozodoi/BMS_Molecular_Translation) with the complete training codes\n\n\n### Data\n- single 80/20 train/test split stratified by the molecule length\n- creating 3M extra images from `extra_approved_InChIs` using RDKit (thanks @tuckerarrants)\n- on each epoch, I was taking all train images + random 1M extra images for training\n\n\n### Tokenizer\n- my tokenizer is very similar to the one proposed by @yasufuminakama in his great pipeline\n- using training + extra data to fit the tokenizer, so it ended up having about 20 more tokens due to more distinct molecules in the data\n\n\n### Image augmentations\n- making lines and letters thicker with cv2 morphological transformations:\n```\nimage = cv2.morphologyEx(image, cv2.MORPH_OPEN, np.ones((2, 2)))\nimage = cv2.erode(image, np.ones((2, 2)))\n```\n- small image rotations during training:\n```\nShiftScaleRotate(0.01, 0.01, 0.10)\n```\n- cropping pictures with large empty borders (thanks @markwijkhuizen)\n- rotating test images with height > width (thanks @hengck23)\n![images](https://i.postimg.cc/t4KqLMNC/inchi.jpg)\n\n\n### Base models\nAll models used EfficientNet-based CNN encoder and LSTM-based RNN decoder. The table below from Neptune.ai provides the main architecture and training parameters of my top models:\n![models](https://i.postimg.cc/cLrTp1Pc/Screen-2021-06-04-at-10-17-02.jpg)\n\n\n### Ensembling & post-processing \n- using beam search with k = 5 on most of the base models (thanks @tugstugi)\n- normalizing each model predictions with RDKit (thanks @nofreewill)\n- ensembling predictions from base models using majority voting\n- setting \"unconfident\" predictions to the ones produced by the lowest-CV model that has a valid prediction according to RDKit (thanks @nofreewill again)\n\n\n### What else I planned to try\n- adding image meta-data such as pixel resolution and height/width ratio as input features\n- iteratively predicting different parts of the InChI (formula, /c and /h parts) with separate models that would take image and predictions of the previous molecule part as input\n- switching to transformers. I only started working on transformer pipeline a few days ago and did not have enough time to train a heavy model due to limited resources\n\nThe final solution achieves **1.32** on the public and **1.31** on the private LB (47th place). I hope this summary was useful for some of you. Happy to answer any questions in the comments and see you in the next competitions! 😊",
      "votes": 31
    },
    {
      "id": 1345865,
      "postDate": "2021-06-11T22:34:52.537Z",
      "content": "<p>Good work, its helpful</p>",
      "rawMarkdown": "Good work, its helpful",
      "votes": 1
    },
    {
      "id": 1344286,
      "postDate": "2021-06-10T19:16:08.737Z",
      "content": "<p>Nice! Good job</p>",
      "rawMarkdown": "Nice! Good job"
    },
    {
      "id": 1339101,
      "postDate": "2021-06-07T03:28:09.053Z",
      "content": "<p>Nice keep up</p>",
      "rawMarkdown": "Nice keep up"
    },
    {
      "id": 1338984,
      "postDate": "2021-06-06T23:24:46.457Z",
      "content": "<p>Thanks for sharing your write-up. I have also read your blog and you have inspired me to try using neptune.ai on my next project. </p>",
      "rawMarkdown": "Thanks for sharing your write-up. I have also read your blog and you have inspired me to try using neptune.ai on my next project. ",
      "replies": [
        {
          "id": 1339404,
          "postDate": "2021-06-07T08:16:31.300Z",
          "content": "<p><a href=\"https://www.kaggle.com/talktocharles\" target=\"_blank\">@talktocharles</a> great to hear that, hope it will be useful for you future projects. Thanks!</p>",
          "rawMarkdown": "@talktocharles great to hear that, hope it will be useful for you future projects. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1338477,
      "postDate": "2021-06-06T13:40:32.973Z",
      "content": "<p>Congrats and thanks for the great write-up and your released code! 👍</p>",
      "rawMarkdown": "Congrats and thanks for the great write-up and your released code! 👍",
      "replies": [
        {
          "id": 1338871,
          "postDate": "2021-06-06T19:52:55.860Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a>! I am glad you found it helpful.</p>",
          "rawMarkdown": "Thanks @alexlwh! I am glad you found it helpful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1344701,
      "postDate": "2021-06-11T04:53:43.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1345865,
      "author_name": "Paramatma Pulivarthi",
      "author_url": "",
      "post_date": "2021-06-11T22:34:52.537000",
      "content": "<p>Good work, its helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1344286,
      "author_name": "Razvan Milicin",
      "author_url": "",
      "post_date": "2021-06-10T19:16:08.737000",
      "content": "<p>Nice! Good job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1339101,
      "author_name": "Mohamed Essam",
      "author_url": "",
      "post_date": "2021-06-07T03:28:09.053000",
      "content": "<p>Nice keep up</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1338984,
      "author_name": "Charles",
      "author_url": "",
      "post_date": "2021-06-06T23:24:46.457000",
      "content": "<p>Thanks for sharing your write-up. I have also read your blog and you have inspired me to try using neptune.ai on my next project. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1339404,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-06-07T08:16:31.300000",
          "content": "<p><a href=\"https://www.kaggle.com/talktocharles\" target=\"_blank\">@talktocharles</a> great to hear that, hope it will be useful for you future projects. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1338477,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-06-06T13:40:32.973000",
      "content": "<p>Congrats and thanks for the great write-up and your released code! 👍</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1338871,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-06-06T19:52:55.860000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a>! I am glad you found it helpful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1344701,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-11T04:53:43.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335454": "### Summary\n\nFirst, let me congratulate all the winners and thank the competition hosts - it was a really interesting challenge! It was great to be a part of a competition requiring both CV and NLP skills and having such a large dataset with a very minimal shakeup. \n\nThe (not so) funny thing is that I was confident that the competition deadline is actually two days later, 05.06. I mistakingly created a calendar entry for the deadline and never checked if it's actually correct. What a surprise for me was to check out the competition page yesterday and realize that it was just 10 hours before the end! 😅\n\nMy solution is an ensemble of seven CNN-LSTM Encoder-Decoder models. All models are implemented in PyTorch and trained on a local machine with Quadro RTX 6000 GPU.\n\n\n### Code\n- [Kaggle notebook](https://www.kaggle.com/kozodoi/47th-place-solution-bms-ensembling) reproducing my submission and ensembling pipeline\n- [GitHub repo](https://github.com/kozodoi/BMS_Molecular_Translation) with the complete training codes\n\n\n### Data\n- single 80/20 train/test split stratified by the molecule length\n- creating 3M extra images from `extra_approved_InChIs` using RDKit (thanks @tuckerarrants)\n- on each epoch, I was taking all train images + random 1M extra images for training\n\n\n### Tokenizer\n- my tokenizer is very similar to the one proposed by @yasufuminakama in his great pipeline\n- using training + extra data to fit the tokenizer, so it ended up having about 20 more tokens due to more distinct molecules in the data\n\n\n### Image augmentations\n- making lines and letters thicker with cv2 morphological transformations:\n```\nimage = cv2.morphologyEx(image, cv2.MORPH_OPEN, np.ones((2, 2)))\nimage = cv2.erode(image, np.ones((2, 2)))\n```\n- small image rotations during training:\n```\nShiftScaleRotate(0.01, 0.01, 0.10)\n```\n- cropping pictures with large empty borders (thanks @markwijkhuizen)\n- rotating test images with height > width (thanks @hengck23)\n![images](https://i.postimg.cc/t4KqLMNC/inchi.jpg)\n\n\n### Base models\nAll models used EfficientNet-based CNN encoder and LSTM-based RNN decoder. The table below from Neptune.ai provides the main architecture and training parameters of my top models:\n![models](https://i.postimg.cc/cLrTp1Pc/Screen-2021-06-04-at-10-17-02.jpg)\n\n\n### Ensembling & post-processing \n- using beam search with k = 5 on most of the base models (thanks @tugstugi)\n- normalizing each model predictions with RDKit (thanks @nofreewill)\n- ensembling predictions from base models using majority voting\n- setting \"unconfident\" predictions to the ones produced by the lowest-CV model that has a valid prediction according to RDKit (thanks @nofreewill again)\n\n\n### What else I planned to try\n- adding image meta-data such as pixel resolution and height/width ratio as input features\n- iteratively predicting different parts of the InChI (formula, /c and /h parts) with separate models that would take image and predictions of the previous molecule part as input\n- switching to transformers. I only started working on transformer pipeline a few days ago and did not have enough time to train a heavy model due to limited resources\n\nThe final solution achieves **1.32** on the public and **1.31** on the private LB (47th place). I hope this summary was useful for some of you. Happy to answer any questions in the comments and see you in the next competitions! 😊",
    "1345865": "Good work, its helpful",
    "1344286": "Nice! Good job",
    "1339101": "Nice keep up",
    "1338984": "Thanks for sharing your write-up. I have also read your blog and you have inspired me to try using neptune.ai on my next project. ",
    "1338477": "Congrats and thanks for the great write-up and your released code! 👍",
    "1344701": ""
  }
}