{
  "id": 243868,
  "title": "70th place solution and code",
  "url": "/competitions/bms-molecular-translation/writeups/2080isenberg-70th-place-solution-and-code",
  "author_name": "",
  "post_date": "2021-06-04T12:21:25.750Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>Thanks to the competition host and for higly useful advices in the discussion! It’s my first competition and I am glad to get a bronze :)</p>\n<p>My solution is very simple and based on smile targets. It was hard to make enough experiments due to I have had only a single 2080ti. To some speed up the training, I used torch amp.</p>\n<p><strong>Summary of the solution</strong></p>\n<p>The main pipeline is similar to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> great <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">notebook</a>, but as encoder I use EfficienNet_B3. The main difference - instead of using InChI-target I use <a href=\"https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system\" target=\"_blank\">Smile</a> notation. On inference I convert Smile back to InChI via <a href=\"https://github.com/kuelumbus/rdkit_platform_wheels\" target=\"_blank\">rdkit</a> and as tokenizer I use <a href=\"https://github.com/XinhaoLi74/SmilesPE\" target=\"_blank\">XinhaoLi74/SmilesPE</a>. Smiles is much more simpler for prediction - shorter notation and less token-classes - and this gave me great improvement in accuracy (from <strong>6.31</strong> to <strong>2.22</strong>).</p>\n<p><strong>Data</strong></p>\n<p>I use syntheic images and extra_approved_InChIs.csv of course - about 1 million images (thanks for the great <a href=\"https://www.kaggle.com/tuckerarrants/inchi-allowed-external-data\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> and <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a>)</p>\n<p>I have make a lot of experiments - in hyperparameters, batchsampler, backbones, etc.,  but not as many as I wanted because training was soooo long. I ended up with 288x288 images size, as a compromise between accuracy and training speed.</p>\n<p>I made an adaptive batchsampler that was quite helpful - it increase probability for a sample to be added in batch by sample's loss during training. So the model was trained on \"harder\" examples, and samples with almost zero loss wasn’t put in batch often.</p>\n<p><strong>Some other notes</strong></p>\n<p>There was some little accuracy improvement (about 1.8 -&gt; 1.6) due to freeze encoder and train LSTM-decored only. It significantly speed up training and allowed to increase batch size in final stage of the training.</p>\n<p>Also as optimizer I use AdamW with weight_decay 0.01 (as <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\" target=\"_blank\">article</a> suggest) but don't sure that it get any influence at all :) As image augmentations were used only gaussian blur and various rotations (by randomly transpose and vertical/horizontal flipping).</p>\n<p><strong>Other experiments I wanted to try</strong></p>\n<ul>\n<li>To split InChI up to 8 indepenpent string layers (separated by the \"/\" notation: \"/b\", \"/t\", \"/m\", \"/s\", etc.) and than train individual models for each layer. Each layer in the InChI describe different information about a molecule, and several models, trained on separate inchi-layers, presumably can get good results.</li>\n<li>To add rotation transoform to any angles (it might help the model to catch some useful feature on images)</li>\n<li>To add beam search and some models ensembling.</li>\n<li>Trying Transformers instead of LSTM</li>\n</ul>\n<p>The source code (pytorch) is available in the github <a href=\"https://github.com/skalinin/BMS-competition\" target=\"_blank\">repository</a>. </p>\n<p>Congrat to all who took a part in the competition:)</p>",
  "messages": [
    {
      "id": "1335609",
      "postDate": "06/04/2021 10:09:01",
      "content": "<p>Hi all!</p>\n<p>Thanks to the competition host and for higly useful advices in the discussion! It’s my first competition and I am glad to get a bronze :)</p>\n<p>My solution is very simple and based on smile targets. It was hard to make enough experiments due to I have had only a single 2080ti. To some speed up the training, I used torch amp.</p>\n<p><strong>Summary of the solution</strong></p>\n<p>The main pipeline is similar to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> great <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">notebook</a>, but as encoder I use EfficienNet_B3. The main difference - instead of using InChI-target I use <a href=\"https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system\" target=\"_blank\">Smile</a> notation. On inference I convert Smile back to InChI via <a href=\"https://github.com/kuelumbus/rdkit_platform_wheels\" target=\"_blank\">rdkit</a> and as tokenizer I use <a href=\"https://github.com/XinhaoLi74/SmilesPE\" target=\"_blank\">XinhaoLi74/SmilesPE</a>. Smiles is much more simpler for prediction - shorter notation and less token-classes - and this gave me great improvement in accuracy (from <strong>6.31</strong> to <strong>2.22</strong>).</p>\n<p><strong>Data</strong></p>\n<p>I use syntheic images and extra_approved_InChIs.csv of course - about 1 million images (thanks for the great <a href=\"https://www.kaggle.com/tuckerarrants/inchi-allowed-external-data\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> and <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a>)</p>\n<p>I have make a lot of experiments - in hyperparameters, batchsampler, backbones, etc.,  but not as many as I wanted because training was soooo long. I ended up with 288x288 images size, as a compromise between accuracy and training speed.</p>\n<p>I made an adaptive batchsampler that was quite helpful - it increase probability for a sample to be added in batch by sample's loss during training. So the model was trained on \"harder\" examples, and samples with almost zero loss wasn’t put in batch often.</p>\n<p><strong>Some other notes</strong></p>\n<p>There was some little accuracy improvement (about 1.8 -&gt; 1.6) due to freeze encoder and train LSTM-decored only. It significantly speed up training and allowed to increase batch size in final stage of the training.</p>\n<p>Also as optimizer I use AdamW with weight_decay 0.01 (as <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\" target=\"_blank\">article</a> suggest) but don't sure that it get any influence at all :) As image augmentations were used only gaussian blur and various rotations (by randomly transpose and vertical/horizontal flipping).</p>\n<p><strong>Other experiments I wanted to try</strong></p>\n<ul>\n<li>To split InChI up to 8 indepenpent string layers (separated by the \"/\" notation: \"/b\", \"/t\", \"/m\", \"/s\", etc.) and than train individual models for each layer. Each layer in the InChI describe different information about a molecule, and several models, trained on separate inchi-layers, presumably can get good results.</li>\n<li>To add rotation transoform to any angles (it might help the model to catch some useful feature on images)</li>\n<li>To add beam search and some models ensembling.</li>\n<li>Trying Transformers instead of LSTM</li>\n</ul>\n<p>The source code (pytorch) is available in the github <a href=\"https://github.com/skalinin/BMS-competition\" target=\"_blank\">repository</a>. </p>\n<p>Congrat to all who took a part in the competition:)</p>",
      "rawMarkdown": "Hi all!\n\nThanks to the competition host and for higly useful advices in the discussion! It’s my first competition and I am glad to get a bronze :)\n\nMy solution is very simple and based on smile targets. It was hard to make enough experiments due to I have had only a single 2080ti. To some speed up the training, I used torch amp.\n\n**Summary of the solution**\n\nThe main pipeline is similar to @yasufuminakama great [notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter), but as encoder I use EfficienNet_B3. The main difference - instead of using InChI-target I use [Smile](https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system) notation. On inference I convert Smile back to InChI via [rdkit](https://github.com/kuelumbus/rdkit_platform_wheels) and as tokenizer I use [XinhaoLi74/SmilesPE](https://github.com/XinhaoLi74/SmilesPE). Smiles is much more simpler for prediction - shorter notation and less token-classes - and this gave me great improvement in accuracy (from **6.31** to **2.22**).\n\n**Data**\n\nI use syntheic images and extra_approved_InChIs.csv of course - about 1 million images (thanks for the great [notebook](https://www.kaggle.com/tuckerarrants/inchi-allowed-external-data) by @tuckerarrants and [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3) by @stainsby)\n\nI have make a lot of experiments - in hyperparameters, batchsampler, backbones, etc.,  but not as many as I wanted because training was soooo long. I ended up with 288x288 images size, as a compromise between accuracy and training speed.\n\nI made an adaptive batchsampler that was quite helpful - it increase probability for a sample to be added in batch by sample's loss during training. So the model was trained on \"harder\" examples, and samples with almost zero loss wasn’t put in batch often.\n\n**Some other notes**\n\nThere was some little accuracy improvement (about 1.8 -> 1.6) due to freeze encoder and train LSTM-decored only. It significantly speed up training and allowed to increase batch size in final stage of the training.\n\nAlso as optimizer I use AdamW with weight_decay 0.01 (as [article](https://www.fast.ai/2018/07/02/adam-weight-decay/) suggest) but don't sure that it get any influence at all :) As image augmentations were used only gaussian blur and various rotations (by randomly transpose and vertical/horizontal flipping).\n\n**Other experiments I wanted to try**\n\n- To split InChI up to 8 indepenpent string layers (separated by the \"/\" notation: \"/b\", \"/t\", \"/m\", \"/s\", etc.) and than train individual models for each layer. Each layer in the InChI describe different information about a molecule, and several models, trained on separate inchi-layers, presumably can get good results.\n- To add rotation transoform to any angles (it might help the model to catch some useful feature on images)\n- To add beam search and some models ensembling.\n- Trying Transformers instead of LSTM\n\n\n\nThe source code (pytorch) is available in the github [repository](https://github.com/skalinin/BMS-competition). \n\nCongrat to all who took a part in the competition:)",
      "votes": null
    },
    {
      "id": "1338791",
      "postDate": "06/06/2021 18:20:06",
      "content": "<p>Thanks for the great write-up and release of your code! 👍</p>",
      "rawMarkdown": "Thanks for the great write-up and release of your code! 👍",
      "votes": null
    },
    {
      "id": "1340381",
      "postDate": "06/07/2021 20:37:46",
      "content": "<p>Awesome solution and github repo, very well organized. Thanks for sharing it!</p>",
      "rawMarkdown": "Awesome solution and github repo, very well organized. Thanks for sharing it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1338791,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/06/2021 18:20:06",
      "content": "<p>Thanks for the great write-up and release of your code! 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340381,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "06/07/2021 20:37:46",
      "content": "<p>Awesome solution and github repo, very well organized. Thanks for sharing it!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335609": "Hi all!\n\nThanks to the competition host and for higly useful advices in the discussion! It’s my first competition and I am glad to get a bronze :)\n\nMy solution is very simple and based on smile targets. It was hard to make enough experiments due to I have had only a single 2080ti. To some speed up the training, I used torch amp.\n\n**Summary of the solution**\n\nThe main pipeline is similar to @yasufuminakama great [notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter), but as encoder I use EfficienNet_B3. The main difference - instead of using InChI-target I use [Smile](https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system) notation. On inference I convert Smile back to InChI via [rdkit](https://github.com/kuelumbus/rdkit_platform_wheels) and as tokenizer I use [XinhaoLi74/SmilesPE](https://github.com/XinhaoLi74/SmilesPE). Smiles is much more simpler for prediction - shorter notation and less token-classes - and this gave me great improvement in accuracy (from **6.31** to **2.22**).\n\n**Data**\n\nI use syntheic images and extra_approved_InChIs.csv of course - about 1 million images (thanks for the great [notebook](https://www.kaggle.com/tuckerarrants/inchi-allowed-external-data) by @tuckerarrants and [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3) by @stainsby)\n\nI have make a lot of experiments - in hyperparameters, batchsampler, backbones, etc.,  but not as many as I wanted because training was soooo long. I ended up with 288x288 images size, as a compromise between accuracy and training speed.\n\nI made an adaptive batchsampler that was quite helpful - it increase probability for a sample to be added in batch by sample's loss during training. So the model was trained on \"harder\" examples, and samples with almost zero loss wasn’t put in batch often.\n\n**Some other notes**\n\nThere was some little accuracy improvement (about 1.8 -> 1.6) due to freeze encoder and train LSTM-decored only. It significantly speed up training and allowed to increase batch size in final stage of the training.\n\nAlso as optimizer I use AdamW with weight_decay 0.01 (as [article](https://www.fast.ai/2018/07/02/adam-weight-decay/) suggest) but don't sure that it get any influence at all :) As image augmentations were used only gaussian blur and various rotations (by randomly transpose and vertical/horizontal flipping).\n\n**Other experiments I wanted to try**\n\n- To split InChI up to 8 indepenpent string layers (separated by the \"/\" notation: \"/b\", \"/t\", \"/m\", \"/s\", etc.) and than train individual models for each layer. Each layer in the InChI describe different information about a molecule, and several models, trained on separate inchi-layers, presumably can get good results.\n- To add rotation transoform to any angles (it might help the model to catch some useful feature on images)\n- To add beam search and some models ensembling.\n- Trying Transformers instead of LSTM\n\n\n\nThe source code (pytorch) is available in the github [repository](https://github.com/skalinin/BMS-competition). \n\nCongrat to all who took a part in the competition:)",
    "1338791": "Thanks for the great write-up and release of your code! 👍",
    "1340381": "Awesome solution and github repo, very well organized. Thanks for sharing it!"
  },
  "source": "meta"
}