{
  "id": 243932,
  "title": "3rd place solution: brief summary",
  "url": "/competitions/bms-molecular-translation/writeups/kyamaro-3rd-place-solution-brief-summary",
  "author_name": "",
  "post_date": "2021-06-06T07:32:55.810Z",
  "votes": 79,
  "comment_count": 24,
  "views": 0,
  "content": "<p>First of all, I would like to thank the competition hosts, and all the participants for their hard work. This competition was very challenging, and I learned a lot.<br>\nI would also like to thank my teammates <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. I am convinced that I could not have achieved this result without them.<br>\nHere I will try to summarize some of the main points of our solution.</p>\n<h1>Overview</h1>\n<p>Our solution consists of three phases:</p>\n<ul>\n<li>Phase1: image captioning training</li>\n<li>Phase2: InChI candidates generation</li>\n<li>Phase3: InChI candidates reranking</li>\n</ul>\n<p><img src=\"https://i.ibb.co/f0PYvhZ/bms-solution-Page-1-2.png\" alt=\"\"></p>\n<h1>Phase1: Image captioning training</h1>\n<p>One important aspect of this competition was the difference in data trends between the training and test data. At first, I struggled with the difference between CV and LB, but then we noticed that the test data had a higher amount of salt and pepper noise. Therefore, we solved the gap between CV and LB by adding augmentation of the salt and pepper noise during training.</p>\n<p><img src=\"https://i.ibb.co/j31fQQB/bms-solution-Page-2-3.png\" alt=\"\"></p>\n<h1>Phase2: InChI candidates generation</h1>\n<p>In the phase3, the likelihood for pairs of an image and an InChI candidate are estimated by multiple models, and the InChI candidate having the highest likelihood is used as the final output. Therefore, in this phase, it is necessary to generate a variety of InChI candidates with good quality.<br>\nWe generated a wide variety of InChI candidates by using various models as shown below and by using beam search:</p>\n<ul>\n<li>Swin Transformer + BERT Decoder (by KF &amp; lyakaap)</li>\n<li>Transformer in Transformer + BERT Decoder (by KF)</li>\n<li>EfficientNet-v2 followed by ViT + BERT Decoder (by lyakaap)</li>\n<li>EfficientNet-B4 + Transformer Decoder (image_size:416x736) (by camaro)</li>\n<li>EfficientNet-B4 + Transformer Encoder/Decoder(image_size:300x600) (by camaro)</li>\n</ul>\n<h1>Phase3: InChI candidates reranking</h1>\n<p>One of the key parts of our solution is reranking. In fact, most of the single model results are around LB0.9~0.8, but we were able to reach the final result (LB0.54) by reranking the generated results of multiple models using the logic shown below.</p>\n<ul>\n<li>rdkit.Chem.MolFromInchi function is used to validate for each InChI candidate. (is_valid)</li>\n<li>For each InChI candidate, calculate the loss (cross entropy / focal loss) used for training in multiple models, and average it across models. (loss)</li>\n<li>Sort the candidates in descending order of “is_valid” → ascending order of “loss”, and the InChI with the highest score is the final output.</li>\n</ul>\n<h1>What didn't work</h1>\n<ul>\n<li>Object detection for atoms and bonds w/ yolov5<ul>\n<li>Reconstruct inchis like Dacon.ai's 1st place solution</li>\n<li>Use detection results as additional input channels</li></ul></li>\n<li>MLM pre-training using extra_approved_InChIs.csv</li>\n<li>Rank learning of InChI candidates (pointwise)<ul>\n<li>Perform real/fake classification and levenshtein’s distance prediction using a fake made by beam search as a negative example.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1335908",
      "postDate": "06/04/2021 13:59:24",
      "content": "<p>First of all, I would like to thank the competition hosts, and all the participants for their hard work. This competition was very challenging, and I learned a lot.<br>\nI would also like to thank my teammates <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. I am convinced that I could not have achieved this result without them.<br>\nHere I will try to summarize some of the main points of our solution.</p>\n<h1>Overview</h1>\n<p>Our solution consists of three phases:</p>\n<ul>\n<li>Phase1: image captioning training</li>\n<li>Phase2: InChI candidates generation</li>\n<li>Phase3: InChI candidates reranking</li>\n</ul>\n<p><img src=\"https://i.ibb.co/f0PYvhZ/bms-solution-Page-1-2.png\" alt=\"\"></p>\n<h1>Phase1: Image captioning training</h1>\n<p>One important aspect of this competition was the difference in data trends between the training and test data. At first, I struggled with the difference between CV and LB, but then we noticed that the test data had a higher amount of salt and pepper noise. Therefore, we solved the gap between CV and LB by adding augmentation of the salt and pepper noise during training.</p>\n<p><img src=\"https://i.ibb.co/j31fQQB/bms-solution-Page-2-3.png\" alt=\"\"></p>\n<h1>Phase2: InChI candidates generation</h1>\n<p>In the phase3, the likelihood for pairs of an image and an InChI candidate are estimated by multiple models, and the InChI candidate having the highest likelihood is used as the final output. Therefore, in this phase, it is necessary to generate a variety of InChI candidates with good quality.<br>\nWe generated a wide variety of InChI candidates by using various models as shown below and by using beam search:</p>\n<ul>\n<li>Swin Transformer + BERT Decoder (by KF &amp; lyakaap)</li>\n<li>Transformer in Transformer + BERT Decoder (by KF)</li>\n<li>EfficientNet-v2 followed by ViT + BERT Decoder (by lyakaap)</li>\n<li>EfficientNet-B4 + Transformer Decoder (image_size:416x736) (by camaro)</li>\n<li>EfficientNet-B4 + Transformer Encoder/Decoder(image_size:300x600) (by camaro)</li>\n</ul>\n<h1>Phase3: InChI candidates reranking</h1>\n<p>One of the key parts of our solution is reranking. In fact, most of the single model results are around LB0.9~0.8, but we were able to reach the final result (LB0.54) by reranking the generated results of multiple models using the logic shown below.</p>\n<ul>\n<li>rdkit.Chem.MolFromInchi function is used to validate for each InChI candidate. (is_valid)</li>\n<li>For each InChI candidate, calculate the loss (cross entropy / focal loss) used for training in multiple models, and average it across models. (loss)</li>\n<li>Sort the candidates in descending order of “is_valid” → ascending order of “loss”, and the InChI with the highest score is the final output.</li>\n</ul>\n<h1>What didn't work</h1>\n<ul>\n<li>Object detection for atoms and bonds w/ yolov5<ul>\n<li>Reconstruct inchis like Dacon.ai's 1st place solution</li>\n<li>Use detection results as additional input channels</li></ul></li>\n<li>MLM pre-training using extra_approved_InChIs.csv</li>\n<li>Rank learning of InChI candidates (pointwise)<ul>\n<li>Perform real/fake classification and levenshtein’s distance prediction using a fake made by beam search as a negative example.</li></ul></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank the competition hosts, and all the participants for their hard work. This competition was very challenging, and I learned a lot.\nI would also like to thank my teammates @lyakaap @bamps53. I am convinced that I could not have achieved this result without them.\nHere I will try to summarize some of the main points of our solution.\n\n# Overview\n\nOur solution consists of three phases:\n\n- Phase1: image captioning training\n- Phase2: InChI candidates generation\n- Phase3: InChI candidates reranking\n\n![](https://i.ibb.co/f0PYvhZ/bms-solution-Page-1-2.png)\n\n# Phase1: Image captioning training\n\nOne important aspect of this competition was the difference in data trends between the training and test data. At first, I struggled with the difference between CV and LB, but then we noticed that the test data had a higher amount of salt and pepper noise. Therefore, we solved the gap between CV and LB by adding augmentation of the salt and pepper noise during training.\n\n![](https://i.ibb.co/j31fQQB/bms-solution-Page-2-3.png)\n\n# Phase2: InChI candidates generation\n\nIn the phase3, the likelihood for pairs of an image and an InChI candidate are estimated by multiple models, and the InChI candidate having the highest likelihood is used as the final output. Therefore, in this phase, it is necessary to generate a variety of InChI candidates with good quality.\nWe generated a wide variety of InChI candidates by using various models as shown below and by using beam search:\n\n- Swin Transformer + BERT Decoder (by KF & lyakaap)\n- Transformer in Transformer + BERT Decoder (by KF)\n- EfficientNet-v2 followed by ViT + BERT Decoder (by lyakaap)\n- EfficientNet-B4 + Transformer Decoder (image_size:416x736) (by camaro)\n- EfficientNet-B4 + Transformer Encoder/Decoder(image_size:300x600) (by camaro)\n\n# Phase3: InChI candidates reranking\n\nOne of the key parts of our solution is reranking. In fact, most of the single model results are around LB0.9~0.8, but we were able to reach the final result (LB0.54) by reranking the generated results of multiple models using the logic shown below.\n\n- rdkit.Chem.MolFromInchi function is used to validate for each InChI candidate. (is_valid)\n- For each InChI candidate, calculate the loss (cross entropy / focal loss) used for training in multiple models, and average it across models. (loss)\n- Sort the candidates in descending order of “is_valid” → ascending order of “loss”, and the InChI with the highest score is the final output.\n\n# What didn't work\n\n- Object detection for atoms and bonds w/ yolov5\n    - Reconstruct inchis like Dacon.ai's 1st place solution\n    - Use detection results as additional input channels\n- MLM pre-training using extra_approved_InChIs.csv\n- Rank learning of InChI candidates (pointwise)\n    - Perform real/fake classification and levenshtein’s distance prediction using a fake made by beam search as a negative example.",
      "votes": null
    },
    {
      "id": "1336027",
      "postDate": "06/04/2021 15:35:32",
      "content": "<p>Congrats for GM!</p>",
      "rawMarkdown": "Congrats for GM!",
      "votes": null
    },
    {
      "id": "1336244",
      "postDate": "06/04/2021 18:26:43",
      "content": "<p>Congrats Grand Master. Can you please give some details on the machine that you used to train your models?</p>",
      "rawMarkdown": "Congrats Grand Master. Can you please give some details on the machine that you used to train your models?",
      "votes": null
    },
    {
      "id": "1336342",
      "postDate": "06/04/2021 19:46:44",
      "content": "<p>Congrats!! Will you share your codes? Thanks!!</p>",
      "rawMarkdown": "Congrats!! Will you share your codes? Thanks!!",
      "votes": null
    },
    {
      "id": "1336373",
      "postDate": "06/04/2021 20:34:35",
      "content": "<p>Congrats GM!</p>",
      "rawMarkdown": "Congrats GM!",
      "votes": null
    },
    {
      "id": "1336450",
      "postDate": "06/04/2021 23:31:59",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": null
    },
    {
      "id": "1336543",
      "postDate": "06/05/2021 03:01:04",
      "content": "<p>Thanks!<br>\nMy best model takes a week with A100 to learn 10epoch.<br>\nSince I mainly used A100 preeptible instance ($0.88/hour) on GCP, it would have cost me about $150 to complete the training of one model 😅</p>",
      "rawMarkdown": "Thanks!\nMy best model takes a week with A100 to learn 10epoch.\nSince I mainly used A100 preeptible instance ($0.88/hour) on GCP, it would have cost me about $150 to complete the training of one model 😅",
      "votes": null
    },
    {
      "id": "1336545",
      "postDate": "06/05/2021 03:03:27",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1336600",
      "postDate": "06/05/2021 04:54:16",
      "content": "<p>Congrats and great solution!<br>\nYou guys found more insight of the dataset than us 👍</p>",
      "rawMarkdown": "Congrats and great solution!\nYou guys found more insight of the dataset than us 👍",
      "votes": null
    },
    {
      "id": "1336623",
      "postDate": "06/05/2021 05:19:21",
      "content": "<p>Thanks for interest!<br>\nCompetition-specific code can be shared, but since it relies on my private experiment management library (a customized version of Catalyst), it might be difficult to share executable code 😏</p>",
      "rawMarkdown": "Thanks for interest!\nCompetition-specific code can be shared, but since it relies on my private experiment management library (a customized version of Catalyst), it might be difficult to share executable code 😏",
      "votes": null
    },
    {
      "id": "1336625",
      "postDate": "06/05/2021 05:20:50",
      "content": "<p>Thanks!<br>\nYour starter notebooks are always helpful 😊</p>",
      "rawMarkdown": "Thanks!\nYour starter notebooks are always helpful 😊",
      "votes": null
    },
    {
      "id": "1336642",
      "postDate": "06/05/2021 05:30:33",
      "content": "<p>Thanks!<br>\nWe knew your team would always come out on top, so we were afraid of your team till the end.<br>\nYour noise injection idea was awesome. I had failed to learn the 12-layer Transformer Decoder (just use 3 layers), but it could have helped.</p>",
      "rawMarkdown": "Thanks!\nWe knew your team would always come out on top, so we were afraid of your team till the end.\nYour noise injection idea was awesome. I had failed to learn the 12-layer Transformer Decoder (just use 3 layers), but it could have helped.",
      "votes": null
    },
    {
      "id": "1336721",
      "postDate": "06/05/2021 07:19:32",
      "content": "<p>Here is my part of solution.</p>\n<h1>TLDR;</h1>\n<p>My best single model achieves lb0.74.<br>\nKey components are below:</p>\n<ul>\n<li>EffientNet B4 + Transformer Decoder is the baseline</li>\n<li>Image size = (416, 736) works best</li>\n<li>Batch size = 128 and 50 epochs training, takes almost 10 days on TPU</li>\n<li>Focal loss is better than cross entropy</li>\n<li>Pseudo labeling</li>\n</ul>\n<p>You can find the code here.  <br>\n<a href=\"https://www.kaggle.com/bamps53/bms-baseline\" target=\"_blank\">https://www.kaggle.com/bamps53/bms-baseline</a></p>\n<hr>\n<h1>Detail</h1>\n<h2>Baseline</h2>\n<p>I've started to work with these great kernels, thanks for sharing!</p>\n<p>Tokenization part is from these great starter kit.  <br>\n<a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a> <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference</a> <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a></p>\n<p>TPU training pipeline is from this kernel, it saves me a lot.  <br>\n<a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92</a></p>\n<p>And transformer part is from here.  <br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-sh\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-sh</a></p>\n<h2>Trainig details</h2>\n<p>After swhiched to transformer from LSTM, I've tried many conbination of patameters.<br>\nAnd then found these insights;</p>\n<ul>\n<li>Image size is most importatnt, bigger is better.</li>\n<li>In my intuition, aspect ratio is useful information, but just resizing to fixed image size works best.</li>\n<li>In encoder, Adding positional encoding to only query and key is better than adding to encoder output directly. It's same as DETR does.</li>\n<li>As my model got to predict very well, most of trainig data got so easy one, so I thought focal loss works here. And it actually did. It was later verified by my teammate after team mergeing.</li>\n<li>After team merging, my teammate shared how to deal with train/test difference, or CV/LB gap. I've approached to the gap by denoise/noise method(shared above) and pseudo labeling. both worked.</li>\n<li>Batch size is 128 and train 50 epochs. It takes to train models about 10 days. Without TPU, it would be a month or more…</li>\n<li>Other small detail can be found in <a href=\"https://www.kaggle.com/bamps53/bms-baseline\" target=\"_blank\">the notebook</a>.</li>\n</ul>\n<p>Here is the final models, we used all model to generate candidate InChI and rescore the candidates.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>Transformer</th>\n<th>Image size</th>\n<th>Add noise</th>\n<th>Denoise</th>\n<th>Pseudo labeling</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>300x600</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.78</td>\n<td>0.96</td>\n</tr>\n<tr>\n<td>2</td>\n<td>EfficientNet B4</td>\n<td>Encoder/Decoder</td>\n<td>300x600</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.76</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>3</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.67</td>\n<td>0.87</td>\n</tr>\n<tr>\n<td>4</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td></td>\n<td></td>\n<td>TRUE</td>\n<td>0.67</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>5</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td>TRUE</td>\n<td></td>\n<td></td>\n<td>0.65</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>6</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td>TRUE</td>\n<td>TRUE</td>\n<td></td>\n<td>0.825</td>\n<td>0.77</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Here is my part of solution.\n\n# TLDR;\nMy best single model achieves lb0.74.\nKey components are below:\n- EffientNet B4 + Transformer Decoder is the baseline\n- Image size = (416, 736) works best\n- Batch size = 128 and 50 epochs training, takes almost 10 days on TPU\n- Focal loss is better than cross entropy\n- Pseudo labeling\n\nYou can find the code here.  \nhttps://www.kaggle.com/bamps53/bms-baseline\n\n---\n# Detail\n\n## Baseline\n\nI've started to work with these great kernels, thanks for sharing!\n\nTokenization part is from these great starter kit.  \nhttps://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\n\nTPU training pipeline is from this kernel, it saves me a lot.  \nhttps://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\n\nAnd transformer part is from here.  \nhttps://www.kaggle.com/aditya08/imagecaptioning-sh\n\n\n## Trainig details\n\nAfter swhiched to transformer from LSTM, I've tried many conbination of patameters.\nAnd then found these insights;\n- Image size is most importatnt, bigger is better.\n- In my intuition, aspect ratio is useful information, but just resizing to fixed image size works best.\n- In encoder, Adding positional encoding to only query and key is better than adding to encoder output directly. It's same as DETR does.\n- As my model got to predict very well, most of trainig data got so easy one, so I thought focal loss works here. And it actually did. It was later verified by my teammate after team mergeing.\n- After team merging, my teammate shared how to deal with train/test difference, or CV/LB gap. I've approached to the gap by denoise/noise method(shared above) and pseudo labeling. both worked.\n- Batch size is 128 and train 50 epochs. It takes to train models about 10 days. Without TPU, it would be a month or more...\n- Other small detail can be found in [the notebook](https://www.kaggle.com/bamps53/bms-baseline).\n\nHere is the final models, we used all model to generate candidate InChI and rescore the candidates.\n\n\n| Model | Backbone        | Transformer     | Image size | Add noise | Denoise | Pseudo labeling | CV    | LB   |\n|-------|-----------------|-----------------|------------|-----------|---------|-----------------|-------|------|\n|     1 | EfficientNet B4 | Decoder only    | 300x600    |           |         |                 |  0.78 | 0.96 |\n|     2 | EfficientNet B4 | Encoder/Decoder | 300x600    |           |         |                 |  0.76 | 0.97 |\n|     3 | EfficientNet B4 | Decoder only    | 416x736    |           |         |                 |  0.67 | 0.87 |\n|     4 | EfficientNet B4 | Decoder only    | 416x736    |           |         |       TRUE      |  0.67 | 0.73 |\n|     5 | EfficientNet B4 | Decoder only    | 416x736    |    TRUE   |         |                 |  0.65 | 0.84 |\n|     6 | EfficientNet B4 | Decoder only    | 416x736    |    TRUE   |   TRUE  |                 | 0.825 | 0.77 |",
      "votes": null
    },
    {
      "id": "1337168",
      "postDate": "06/05/2021 13:05:45",
      "content": "<p>thanks for sharing.  I planned to use google cloud, however after reading docs, I am not sure about pricing detail and also got lost in the platform interface, too much stuff, I was scared away.</p>",
      "rawMarkdown": "thanks for sharing.  I planned to use google cloud, however after reading docs, I am not sure about pricing detail and also got lost in the platform interface, too much stuff, I was scared away.",
      "votes": null
    },
    {
      "id": "1337518",
      "postDate": "06/05/2021 16:59:33",
      "content": "<p>Congrats for GM!<br>\nIts a great write-up and ur diagram looks awesome!<br>\nDid u use any special app to draw them?</p>",
      "rawMarkdown": "Congrats for GM!\nIts a great write-up and ur diagram looks awesome!\nDid u use any special app to draw them?",
      "votes": null
    },
    {
      "id": "1338024",
      "postDate": "06/06/2021 05:43:08",
      "content": "<p>Thanks!<br>\nI used diagrams.net with google drive:<br>\n<a href=\"https://www.diagrams.net/doc/faq/google-drive-diagrams\" target=\"_blank\">https://www.diagrams.net/doc/faq/google-drive-diagrams</a></p>",
      "rawMarkdown": "Thanks!\nI used diagrams.net with google drive:\nhttps://www.diagrams.net/doc/faq/google-drive-diagrams",
      "votes": null
    },
    {
      "id": "1338039",
      "postDate": "06/06/2021 05:50:27",
      "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a>  <br>\nthanks for sharing</p>",
      "rawMarkdown": "kfujikawa  \nthanks for sharing",
      "votes": null
    },
    {
      "id": "1339157",
      "postDate": "06/07/2021 04:55:24",
      "content": "<p>Congratulations to <a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a>, <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> and <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>! Very well deserved!</p>",
      "rawMarkdown": "Congratulations to @kfujikawa, @lyakaap and @bamps53! Very well deserved!",
      "votes": null
    },
    {
      "id": "1339716",
      "postDate": "06/07/2021 12:20:58",
      "content": "<p><a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> Thanks, please give congrats to <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> too😂</p>",
      "rawMarkdown": "saurabhbagchi Thanks, please give congrats to @lyakaap too😂",
      "votes": null
    },
    {
      "id": "1339867",
      "postDate": "06/07/2021 13:53:57",
      "content": "<p>Done, thanks for the correction, update previous comment!</p>",
      "rawMarkdown": "Done, thanks for the correction, update previous comment!",
      "votes": null
    },
    {
      "id": "1340778",
      "postDate": "06/08/2021 08:57:44",
      "content": "<p>Congratulations GM. You deserved it.</p>",
      "rawMarkdown": "Congratulations GM. You deserved it.",
      "votes": null
    },
    {
      "id": "1355271",
      "postDate": "06/18/2021 07:44:43",
      "content": "<p>Congrats and great solution!<br>\nCan share the code: Transformer Encoder/Decoder with beam search . Thanks!</p>",
      "rawMarkdown": "Congrats and great solution!\nCan share the code: Transformer Encoder/Decoder with beam search . Thanks!",
      "votes": null
    },
    {
      "id": "1390636",
      "postDate": "07/16/2021 22:04:48",
      "content": "<p>Thank you for your solution! In your code, after I train a model 'trainer', how can I run it to output a prediction for a custom image if I have a string path_to_file, e.g. '/kaggle/input/dataset/image1.png'? I tried using tf.keras.preprocessing.image_dataset_from_directory but it keeps throwing errors</p>",
      "rawMarkdown": "Thank you for your solution! In your code, after I train a model 'trainer', how can I run it to output a prediction for a custom image if I have a string path_to_file, e.g. '/kaggle/input/dataset/image1.png'? I tried using tf.keras.preprocessing.image_dataset_from_directory but it keeps throwing errors",
      "votes": null
    },
    {
      "id": "1536474",
      "postDate": "10/06/2021 18:04:40",
      "content": "<p><a href=\"https://www.kaggle.com/KF\" target=\"_blank\">@KF</a> Thanks for sharing! Could you please point me to article about BERT decoder you used?</p>",
      "rawMarkdown": "KF Thanks for sharing! Could you please point me to article about BERT decoder you used?",
      "votes": null
    },
    {
      "id": "1536881",
      "postDate": "10/07/2021 05:26:43",
      "content": "<p>I used <a href=\"https://github.com/huggingface/transformers/blob/27d4639779d2d316a7c5f18d22f22d2565b84e5e/src/transformers/models/bert/modeling_bert.py#L1131\" target=\"_blank\">BertLMHeadModel</a> as a decoder.</p>",
      "rawMarkdown": "I used [BertLMHeadModel](https://github.com/huggingface/transformers/blob/27d4639779d2d316a7c5f18d22f22d2565b84e5e/src/transformers/models/bert/modeling_bert.py#L1131) as a decoder.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1336027,
      "author_name": "akirasosa",
      "author_url": "",
      "post_date": "06/04/2021 15:35:32",
      "content": "<p>Congrats for GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336545,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/05/2021 03:03:27",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336244,
      "author_name": "vikrant06",
      "author_url": "",
      "post_date": "06/04/2021 18:26:43",
      "content": "<p>Congrats Grand Master. Can you please give some details on the machine that you used to train your models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336543,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/05/2021 03:01:04",
          "content": "<p>Thanks!<br>\nMy best model takes a week with A100 to learn 10epoch.<br>\nSince I mainly used A100 preeptible instance ($0.88/hour) on GCP, it would have cost me about $150 to complete the training of one model 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337168,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "06/05/2021 13:05:45",
          "content": "<p>thanks for sharing.  I planned to use google cloud, however after reading docs, I am not sure about pricing detail and also got lost in the platform interface, too much stuff, I was scared away.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336342,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "06/04/2021 19:46:44",
      "content": "<p>Congrats!! Will you share your codes? Thanks!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336623,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/05/2021 05:19:21",
          "content": "<p>Thanks for interest!<br>\nCompetition-specific code can be shared, but since it relies on my private experiment management library (a customized version of Catalyst), it might be difficult to share executable code 😏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336373,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "06/04/2021 20:34:35",
      "content": "<p>Congrats GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336625,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/05/2021 05:20:50",
          "content": "<p>Thanks!<br>\nYour starter notebooks are always helpful 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336450,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/04/2021 23:31:59",
      "content": "<p>thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336600,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "06/05/2021 04:54:16",
      "content": "<p>Congrats and great solution!<br>\nYou guys found more insight of the dataset than us 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336642,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/05/2021 05:30:33",
          "content": "<p>Thanks!<br>\nWe knew your team would always come out on top, so we were afraid of your team till the end.<br>\nYour noise injection idea was awesome. I had failed to learn the 12-layer Transformer Decoder (just use 3 layers), but it could have helped.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336721,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "06/05/2021 07:19:32",
      "content": "<p>Here is my part of solution.</p>\n<h1>TLDR;</h1>\n<p>My best single model achieves lb0.74.<br>\nKey components are below:</p>\n<ul>\n<li>EffientNet B4 + Transformer Decoder is the baseline</li>\n<li>Image size = (416, 736) works best</li>\n<li>Batch size = 128 and 50 epochs training, takes almost 10 days on TPU</li>\n<li>Focal loss is better than cross entropy</li>\n<li>Pseudo labeling</li>\n</ul>\n<p>You can find the code here.  <br>\n<a href=\"https://www.kaggle.com/bamps53/bms-baseline\" target=\"_blank\">https://www.kaggle.com/bamps53/bms-baseline</a></p>\n<hr>\n<h1>Detail</h1>\n<h2>Baseline</h2>\n<p>I've started to work with these great kernels, thanks for sharing!</p>\n<p>Tokenization part is from these great starter kit.  <br>\n<a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a> <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference</a> <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a></p>\n<p>TPU training pipeline is from this kernel, it saves me a lot.  <br>\n<a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92</a></p>\n<p>And transformer part is from here.  <br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-sh\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-sh</a></p>\n<h2>Trainig details</h2>\n<p>After swhiched to transformer from LSTM, I've tried many conbination of patameters.<br>\nAnd then found these insights;</p>\n<ul>\n<li>Image size is most importatnt, bigger is better.</li>\n<li>In my intuition, aspect ratio is useful information, but just resizing to fixed image size works best.</li>\n<li>In encoder, Adding positional encoding to only query and key is better than adding to encoder output directly. It's same as DETR does.</li>\n<li>As my model got to predict very well, most of trainig data got so easy one, so I thought focal loss works here. And it actually did. It was later verified by my teammate after team mergeing.</li>\n<li>After team merging, my teammate shared how to deal with train/test difference, or CV/LB gap. I've approached to the gap by denoise/noise method(shared above) and pseudo labeling. both worked.</li>\n<li>Batch size is 128 and train 50 epochs. It takes to train models about 10 days. Without TPU, it would be a month or more…</li>\n<li>Other small detail can be found in <a href=\"https://www.kaggle.com/bamps53/bms-baseline\" target=\"_blank\">the notebook</a>.</li>\n</ul>\n<p>Here is the final models, we used all model to generate candidate InChI and rescore the candidates.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>Transformer</th>\n<th>Image size</th>\n<th>Add noise</th>\n<th>Denoise</th>\n<th>Pseudo labeling</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>300x600</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.78</td>\n<td>0.96</td>\n</tr>\n<tr>\n<td>2</td>\n<td>EfficientNet B4</td>\n<td>Encoder/Decoder</td>\n<td>300x600</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.76</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>3</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.67</td>\n<td>0.87</td>\n</tr>\n<tr>\n<td>4</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td></td>\n<td></td>\n<td>TRUE</td>\n<td>0.67</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>5</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td>TRUE</td>\n<td></td>\n<td></td>\n<td>0.65</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>6</td>\n<td>EfficientNet B4</td>\n<td>Decoder only</td>\n<td>416x736</td>\n<td>TRUE</td>\n<td>TRUE</td>\n<td></td>\n<td>0.825</td>\n<td>0.77</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 1390636,
          "author_name": "rizorakhmanov",
          "author_url": "",
          "post_date": "07/16/2021 22:04:48",
          "content": "<p>Thank you for your solution! In your code, after I train a model 'trainer', how can I run it to output a prediction for a custom image if I have a string path_to_file, e.g. '/kaggle/input/dataset/image1.png'? I tried using tf.keras.preprocessing.image_dataset_from_directory but it keeps throwing errors</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337518,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/05/2021 16:59:33",
      "content": "<p>Congrats for GM!<br>\nIts a great write-up and ur diagram looks awesome!<br>\nDid u use any special app to draw them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1338024,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "06/06/2021 05:43:08",
          "content": "<p>Thanks!<br>\nI used diagrams.net with google drive:<br>\n<a href=\"https://www.diagrams.net/doc/faq/google-drive-diagrams\" target=\"_blank\">https://www.diagrams.net/doc/faq/google-drive-diagrams</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338039,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "06/06/2021 05:50:27",
          "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a>  <br>\nthanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1339157,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "06/07/2021 04:55:24",
      "content": "<p>Congratulations to <a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a>, <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> and <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>! Very well deserved!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1339716,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "06/07/2021 12:20:58",
          "content": "<p><a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> Thanks, please give congrats to <a href=\"https://www.kaggle.com/lyakaap\" target=\"_blank\">@lyakaap</a> too😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1339867,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "06/07/2021 13:53:57",
          "content": "<p>Done, thanks for the correction, update previous comment!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1340778,
      "author_name": "sohailds",
      "author_url": "",
      "post_date": "06/08/2021 08:57:44",
      "content": "<p>Congratulations GM. You deserved it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1355271,
      "author_name": "dragonfirea",
      "author_url": "",
      "post_date": "06/18/2021 07:44:43",
      "content": "<p>Congrats and great solution!<br>\nCan share the code: Transformer Encoder/Decoder with beam search . Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1536474,
      "author_name": "rednikotin",
      "author_url": "",
      "post_date": "10/06/2021 18:04:40",
      "content": "<p><a href=\"https://www.kaggle.com/KF\" target=\"_blank\">@KF</a> Thanks for sharing! Could you please point me to article about BERT decoder you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1536881,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "10/07/2021 05:26:43",
          "content": "<p>I used <a href=\"https://github.com/huggingface/transformers/blob/27d4639779d2d316a7c5f18d22f22d2565b84e5e/src/transformers/models/bert/modeling_bert.py#L1131\" target=\"_blank\">BertLMHeadModel</a> as a decoder.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1335908": "First of all, I would like to thank the competition hosts, and all the participants for their hard work. This competition was very challenging, and I learned a lot.\nI would also like to thank my teammates @lyakaap @bamps53. I am convinced that I could not have achieved this result without them.\nHere I will try to summarize some of the main points of our solution.\n\n# Overview\n\nOur solution consists of three phases:\n\n- Phase1: image captioning training\n- Phase2: InChI candidates generation\n- Phase3: InChI candidates reranking\n\n![](https://i.ibb.co/f0PYvhZ/bms-solution-Page-1-2.png)\n\n# Phase1: Image captioning training\n\nOne important aspect of this competition was the difference in data trends between the training and test data. At first, I struggled with the difference between CV and LB, but then we noticed that the test data had a higher amount of salt and pepper noise. Therefore, we solved the gap between CV and LB by adding augmentation of the salt and pepper noise during training.\n\n![](https://i.ibb.co/j31fQQB/bms-solution-Page-2-3.png)\n\n# Phase2: InChI candidates generation\n\nIn the phase3, the likelihood for pairs of an image and an InChI candidate are estimated by multiple models, and the InChI candidate having the highest likelihood is used as the final output. Therefore, in this phase, it is necessary to generate a variety of InChI candidates with good quality.\nWe generated a wide variety of InChI candidates by using various models as shown below and by using beam search:\n\n- Swin Transformer + BERT Decoder (by KF & lyakaap)\n- Transformer in Transformer + BERT Decoder (by KF)\n- EfficientNet-v2 followed by ViT + BERT Decoder (by lyakaap)\n- EfficientNet-B4 + Transformer Decoder (image_size:416x736) (by camaro)\n- EfficientNet-B4 + Transformer Encoder/Decoder(image_size:300x600) (by camaro)\n\n# Phase3: InChI candidates reranking\n\nOne of the key parts of our solution is reranking. In fact, most of the single model results are around LB0.9~0.8, but we were able to reach the final result (LB0.54) by reranking the generated results of multiple models using the logic shown below.\n\n- rdkit.Chem.MolFromInchi function is used to validate for each InChI candidate. (is_valid)\n- For each InChI candidate, calculate the loss (cross entropy / focal loss) used for training in multiple models, and average it across models. (loss)\n- Sort the candidates in descending order of “is_valid” → ascending order of “loss”, and the InChI with the highest score is the final output.\n\n# What didn't work\n\n- Object detection for atoms and bonds w/ yolov5\n    - Reconstruct inchis like Dacon.ai's 1st place solution\n    - Use detection results as additional input channels\n- MLM pre-training using extra_approved_InChIs.csv\n- Rank learning of InChI candidates (pointwise)\n    - Perform real/fake classification and levenshtein’s distance prediction using a fake made by beam search as a negative example.",
    "1336027": "Congrats for GM!",
    "1336244": "Congrats Grand Master. Can you please give some details on the machine that you used to train your models?",
    "1336342": "Congrats!! Will you share your codes? Thanks!!",
    "1336373": "Congrats GM!",
    "1336450": "thanks for sharing!",
    "1336543": "Thanks!\nMy best model takes a week with A100 to learn 10epoch.\nSince I mainly used A100 preeptible instance ($0.88/hour) on GCP, it would have cost me about $150 to complete the training of one model 😅",
    "1336545": "Thank you!",
    "1336600": "Congrats and great solution!\nYou guys found more insight of the dataset than us 👍",
    "1336623": "Thanks for interest!\nCompetition-specific code can be shared, but since it relies on my private experiment management library (a customized version of Catalyst), it might be difficult to share executable code 😏",
    "1336625": "Thanks!\nYour starter notebooks are always helpful 😊",
    "1336642": "Thanks!\nWe knew your team would always come out on top, so we were afraid of your team till the end.\nYour noise injection idea was awesome. I had failed to learn the 12-layer Transformer Decoder (just use 3 layers), but it could have helped.",
    "1336721": "Here is my part of solution.\n\n# TLDR;\nMy best single model achieves lb0.74.\nKey components are below:\n- EffientNet B4 + Transformer Decoder is the baseline\n- Image size = (416, 736) works best\n- Batch size = 128 and 50 epochs training, takes almost 10 days on TPU\n- Focal loss is better than cross entropy\n- Pseudo labeling\n\nYou can find the code here.  \nhttps://www.kaggle.com/bamps53/bms-baseline\n\n---\n# Detail\n\n## Baseline\n\nI've started to work with these great kernels, thanks for sharing!\n\nTokenization part is from these great starter kit.  \nhttps://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\n\nTPU training pipeline is from this kernel, it saves me a lot.  \nhttps://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\n\nAnd transformer part is from here.  \nhttps://www.kaggle.com/aditya08/imagecaptioning-sh\n\n\n## Trainig details\n\nAfter swhiched to transformer from LSTM, I've tried many conbination of patameters.\nAnd then found these insights;\n- Image size is most importatnt, bigger is better.\n- In my intuition, aspect ratio is useful information, but just resizing to fixed image size works best.\n- In encoder, Adding positional encoding to only query and key is better than adding to encoder output directly. It's same as DETR does.\n- As my model got to predict very well, most of trainig data got so easy one, so I thought focal loss works here. And it actually did. It was later verified by my teammate after team mergeing.\n- After team merging, my teammate shared how to deal with train/test difference, or CV/LB gap. I've approached to the gap by denoise/noise method(shared above) and pseudo labeling. both worked.\n- Batch size is 128 and train 50 epochs. It takes to train models about 10 days. Without TPU, it would be a month or more...\n- Other small detail can be found in [the notebook](https://www.kaggle.com/bamps53/bms-baseline).\n\nHere is the final models, we used all model to generate candidate InChI and rescore the candidates.\n\n\n| Model | Backbone        | Transformer     | Image size | Add noise | Denoise | Pseudo labeling | CV    | LB   |\n|-------|-----------------|-----------------|------------|-----------|---------|-----------------|-------|------|\n|     1 | EfficientNet B4 | Decoder only    | 300x600    |           |         |                 |  0.78 | 0.96 |\n|     2 | EfficientNet B4 | Encoder/Decoder | 300x600    |           |         |                 |  0.76 | 0.97 |\n|     3 | EfficientNet B4 | Decoder only    | 416x736    |           |         |                 |  0.67 | 0.87 |\n|     4 | EfficientNet B4 | Decoder only    | 416x736    |           |         |       TRUE      |  0.67 | 0.73 |\n|     5 | EfficientNet B4 | Decoder only    | 416x736    |    TRUE   |         |                 |  0.65 | 0.84 |\n|     6 | EfficientNet B4 | Decoder only    | 416x736    |    TRUE   |   TRUE  |                 | 0.825 | 0.77 |",
    "1337168": "thanks for sharing.  I planned to use google cloud, however after reading docs, I am not sure about pricing detail and also got lost in the platform interface, too much stuff, I was scared away.",
    "1337518": "Congrats for GM!\nIts a great write-up and ur diagram looks awesome!\nDid u use any special app to draw them?",
    "1338024": "Thanks!\nI used diagrams.net with google drive:\nhttps://www.diagrams.net/doc/faq/google-drive-diagrams",
    "1338039": "kfujikawa  \nthanks for sharing",
    "1339157": "Congratulations to @kfujikawa, @lyakaap and @bamps53! Very well deserved!",
    "1339716": "saurabhbagchi Thanks, please give congrats to @lyakaap too😂",
    "1339867": "Done, thanks for the correction, update previous comment!",
    "1340778": "Congratulations GM. You deserved it.",
    "1355271": "Congrats and great solution!\nCan share the code: Transformer Encoder/Decoder with beam search . Thanks!",
    "1390636": "Thank you for your solution! In your code, after I train a model 'trainer', how can I run it to output a prediction for a custom image if I have a string path_to_file, e.g. '/kaggle/input/dataset/image1.png'? I tried using tf.keras.preprocessing.image_dataset_from_directory but it keeps throwing errors",
    "1536474": "KF Thanks for sharing! Could you please point me to article about BERT decoder you used?",
    "1536881": "I used [BertLMHeadModel](https://github.com/huggingface/transformers/blob/27d4639779d2d316a7c5f18d22f22d2565b84e5e/src/transformers/models/bert/modeling_bert.py#L1131) as a decoder."
  },
  "source": "meta"
}