{
  "id": 244078,
  "title": "8th Place Solution🔺",
  "url": "/competitions/bms-molecular-translation/writeups/charm-8th-place-solution",
  "author_name": "",
  "post_date": "2021-06-05T12:01:07.977Z",
  "votes": 36,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Congratulations to all the winners and thank the competition hosts for holding such a wonderful competition.<br>\nAnd especially thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing the great ideas and code. I think his post made this competition very active.</p>\n<p>This competition was very tough for a solo kaggler like me and I was very inspired by other solo kagglers ( <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>, <a href=\"https://www.kaggle.com/AkiraSosa\" target=\"_blank\">@AkiraSosa</a>, <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>) on LB. Finally, I got solo gold and was able to become a Grandmaster at the expense of health and many holidays!</p>\n<h1>Models</h1>\n<p>I chose two models to train, SwinTransformer with image size 384x384 and patchwise TNT and I implemented it by referring to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s codes. <br>\nI used hyper parameter embed_dim=768, in_dim=48 for TNT and this is larger than TNT-B.</p>\n<p>SwinTransformer and TNT both achieved LB~0.8 as single models and their CV were LD=0.80/0.80, CELoss=0.0030/0.0020</p>\n<h1>Data</h1>\n<p>I generated about 10M images using all the InChIs in <code>extra_approved_InChIs.csv</code> using <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a>'s <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a>. About 13M images are used in training. In the early stage of training, the loss for extra images was bad, but after long training, the loss for extra images and valid score improved. I trained SwinTransformer 15 epochs and TNT 20 epochs. SwinTransformer took 20 hours for 1 epoch and TNT took 12 hours for 1 epoch with 8 x V100.</p>\n<h1>Ensemble</h1>\n<p>I ensembled 8 models in each decoder step (weighted average of predict_proba). Models are trained competition data+extra data or finetuned by only competition data after that or different fold. It seems important to set total weight of SwinTransformer and TNT equally. CV score of ensemble model was LD=0.60 and CELoss=0.0013 and LB:0.63</p>\n<h1>Beam Search</h1>\n<p>On the last day of the competition, I implemented beam search by referring to fairseq <a href=\"https://github.com/pytorch/fairseq/blob/master/fairseq/sequence_generator.py\" target=\"_blank\">code</a>. <br>\nBy performing beam search with beam size 16, I could get LB:0.60.<br>\nIn the later stages of the competition, I found that the prediction discrepancies of the different models of test data were only about 20k. Therefore, beam search or ensemble inference was performed only on 20k images.</p>\n<h1>PostProcess</h1>\n<p>I always used the rdkit validation postprocess shared by <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>.<br>\nPost-processing improvement diminished with score improvement, but a little improvement remained even at the end.</p>\n<h1>What didn't work</h1>\n<ul>\n<li>I trained a generative model by using InChIs from the  competition training data and generated 8M images(molecules), but it did not help improve CV.</li>\n<li>I trained the object detection model to detect atoms and bonds like the DACON solution and overlaid bboxes on the images, but it did not help improve CV.</li>\n<li>TTA(rotation, flip)</li>\n</ul>",
  "messages": [
    {
      "id": "1336729",
      "postDate": "06/05/2021 07:24:50",
      "content": "<p>Congratulations to all the winners and thank the competition hosts for holding such a wonderful competition.<br>\nAnd especially thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing the great ideas and code. I think his post made this competition very active.</p>\n<p>This competition was very tough for a solo kaggler like me and I was very inspired by other solo kagglers ( <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>, <a href=\"https://www.kaggle.com/AkiraSosa\" target=\"_blank\">@AkiraSosa</a>, <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>) on LB. Finally, I got solo gold and was able to become a Grandmaster at the expense of health and many holidays!</p>\n<h1>Models</h1>\n<p>I chose two models to train, SwinTransformer with image size 384x384 and patchwise TNT and I implemented it by referring to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s codes. <br>\nI used hyper parameter embed_dim=768, in_dim=48 for TNT and this is larger than TNT-B.</p>\n<p>SwinTransformer and TNT both achieved LB~0.8 as single models and their CV were LD=0.80/0.80, CELoss=0.0030/0.0020</p>\n<h1>Data</h1>\n<p>I generated about 10M images using all the InChIs in <code>extra_approved_InChIs.csv</code> using <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a>'s <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a>. About 13M images are used in training. In the early stage of training, the loss for extra images was bad, but after long training, the loss for extra images and valid score improved. I trained SwinTransformer 15 epochs and TNT 20 epochs. SwinTransformer took 20 hours for 1 epoch and TNT took 12 hours for 1 epoch with 8 x V100.</p>\n<h1>Ensemble</h1>\n<p>I ensembled 8 models in each decoder step (weighted average of predict_proba). Models are trained competition data+extra data or finetuned by only competition data after that or different fold. It seems important to set total weight of SwinTransformer and TNT equally. CV score of ensemble model was LD=0.60 and CELoss=0.0013 and LB:0.63</p>\n<h1>Beam Search</h1>\n<p>On the last day of the competition, I implemented beam search by referring to fairseq <a href=\"https://github.com/pytorch/fairseq/blob/master/fairseq/sequence_generator.py\" target=\"_blank\">code</a>. <br>\nBy performing beam search with beam size 16, I could get LB:0.60.<br>\nIn the later stages of the competition, I found that the prediction discrepancies of the different models of test data were only about 20k. Therefore, beam search or ensemble inference was performed only on 20k images.</p>\n<h1>PostProcess</h1>\n<p>I always used the rdkit validation postprocess shared by <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>.<br>\nPost-processing improvement diminished with score improvement, but a little improvement remained even at the end.</p>\n<h1>What didn't work</h1>\n<ul>\n<li>I trained a generative model by using InChIs from the  competition training data and generated 8M images(molecules), but it did not help improve CV.</li>\n<li>I trained the object detection model to detect atoms and bonds like the DACON solution and overlaid bboxes on the images, but it did not help improve CV.</li>\n<li>TTA(rotation, flip)</li>\n</ul>",
      "rawMarkdown": "Congratulations to all the winners and thank the competition hosts for holding such a wonderful competition.\nAnd especially thanks to @hengck23 for sharing the great ideas and code. I think his post made this competition very active.\n\nThis competition was very tough for a solo kaggler like me and I was very inspired by other solo kagglers ( @fergusoci, @AkiraSosa, @nofreewill) on LB. Finally, I got solo gold and was able to become a Grandmaster at the expense of health and many holidays!\n\n\n# Models\nI chose two models to train, SwinTransformer with image size 384x384 and patchwise TNT and I implemented it by referring to @hengck23's codes. \nI used hyper parameter embed_dim=768, in_dim=48 for TNT and this is larger than TNT-B.\n\nSwinTransformer and TNT both achieved LB~0.8 as single models and their CV were LD=0.80/0.80, CELoss=0.0030/0.0020\n\n\n# Data\nI generated about 10M images using all the InChIs in `extra_approved_InChIs.csv` using @stainsby's [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3). About 13M images are used in training. In the early stage of training, the loss for extra images was bad, but after long training, the loss for extra images and valid score improved. I trained SwinTransformer 15 epochs and TNT 20 epochs. SwinTransformer took 20 hours for 1 epoch and TNT took 12 hours for 1 epoch with 8 x V100.\n\n\n# Ensemble\nI ensembled 8 models in each decoder step (weighted average of predict_proba). Models are trained competition data+extra data or finetuned by only competition data after that or different fold. It seems important to set total weight of SwinTransformer and TNT equally. CV score of ensemble model was LD=0.60 and CELoss=0.0013 and LB:0.63\n\n\n# Beam Search\nOn the last day of the competition, I implemented beam search by referring to fairseq [code](https://github.com/pytorch/fairseq/blob/master/fairseq/sequence_generator.py). \nBy performing beam search with beam size 16, I could get LB:0.60.\nIn the later stages of the competition, I found that the prediction discrepancies of the different models of test data were only about 20k. Therefore, beam search or ensemble inference was performed only on 20k images.\n\n\n# PostProcess\nI always used the rdkit validation postprocess shared by @nofreewill.\nPost-processing improvement diminished with score improvement, but a little improvement remained even at the end.\n\n\n# What didn't work\n- I trained a generative model by using InChIs from the  competition training data and generated 8M images(molecules), but it did not help improve CV.\n- I trained the object detection model to detect atoms and bonds like the DACON solution and overlaid bboxes on the images, but it did not help improve CV.\n- TTA(rotation, flip)",
      "votes": null
    },
    {
      "id": "1336809",
      "postDate": "06/05/2021 08:17:34",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>! Delighted that you made it to GM with this one!</p>",
      "rawMarkdown": "Congratulations, @charmq! Delighted that you made it to GM with this one!",
      "votes": null
    },
    {
      "id": "1336903",
      "postDate": "06/05/2021 09:22:51",
      "content": "<p>Very nice! And congrats again for becoming GM. Well deserved:)</p>",
      "rawMarkdown": "Very nice! And congrats again for becoming GM. Well deserved:)",
      "votes": null
    },
    {
      "id": "1336936",
      "postDate": "06/05/2021 09:52:27",
      "content": "<p>Thanks for the great write-up and congrats for the solo gold and becoming GM! </p>",
      "rawMarkdown": "Thanks for the great write-up and congrats for the solo gold and becoming GM!",
      "votes": null
    },
    {
      "id": "1336951",
      "postDate": "06/05/2021 10:02:09",
      "content": "<p>Congratulation for becoming GM. I have big respect for solo teams in top 20.</p>",
      "rawMarkdown": "Congratulation for becoming GM. I have big respect for solo teams in top 20.",
      "votes": null
    },
    {
      "id": "1336976",
      "postDate": "06/05/2021 10:32:19",
      "content": "<p>\"On the last day of the competition, I implemented beam search by referring to fairseq code.\"</p>\n<p>Congratulations on your good work!<br>\ncan you open-source k-beam search code?</p>",
      "rawMarkdown": "\"On the last day of the competition, I implemented beam search by referring to fairseq code.\"\n\nCongratulations on your good work!\ncan you open-source k-beam search code?",
      "votes": null
    },
    {
      "id": "1337016",
      "postDate": "06/05/2021 10:58:15",
      "content": "<p>Congrats! You deserve it. Get some rest and feel better :)</p>",
      "rawMarkdown": "Congrats! You deserve it. Get some rest and feel better :)",
      "votes": null
    },
    {
      "id": "1337072",
      "postDate": "06/05/2021 11:55:44",
      "content": "<p>Thank you for your comment!<br>\nMy code is too dirty to share, but the beam search implementation is mostly a copy of fairseq repo and I think what you should implement yourself is reordering incremental states part. It is very easy to implement reordering like this.</p>\n<pre><code>def reorder_incremental_state(self, incremental_state, new_order):\n   for layer in self.text_decode.layer:\n      layer.self_attn.reorder_incremental_state(incremental_state, new_order)\n      layer.encoder_attn.reorder_incremental_state(incremental_state, new_order)\n</code></pre>",
      "rawMarkdown": "Thank you for your comment!\nMy code is too dirty to share, but the beam search implementation is mostly a copy of fairseq repo and I think what you should implement yourself is reordering incremental states part. It is very easy to implement reordering like this.\n\n~~~\ndef reorder_incremental_state(self, incremental_state, new_order):\n   for layer in self.text_decode.layer:\n      layer.self_attn.reorder_incremental_state(incremental_state, new_order)\n      layer.encoder_attn.reorder_incremental_state(incremental_state, new_order)\n~~~",
      "votes": null
    },
    {
      "id": "1337100",
      "postDate": "06/05/2021 12:16:31",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for solo gold and become GM! Thanks for sharing solution ;)</p>",
      "rawMarkdown": "Congrats @charmq for solo gold and become GM! Thanks for sharing solution ;)",
      "votes": null
    },
    {
      "id": "1337117",
      "postDate": "06/05/2021 12:31:40",
      "content": "<p>Congrats GM!</p>",
      "rawMarkdown": "Congrats GM!",
      "votes": null
    },
    {
      "id": "1337152",
      "postDate": "06/05/2021 12:52:40",
      "content": "<p>Thanks for sharing.<br>\nbeside rdkit validation postprocess, what other process did you use to reduce LB/CV gap?  most other toppers use both rdkit validation and  multiple model voting etc. to correct invalid prediction.</p>\n<p>My LB/CV are 2.35/1.1.   Just wonder rdkit validation itself can reduce how much gap?</p>",
      "rawMarkdown": "Thanks for sharing.\nbeside rdkit validation postprocess, what other process did you use to reduce LB/CV gap?  most other toppers use both rdkit validation and  multiple model voting etc. to correct invalid prediction.\n\nMy LB/CV are 2.35/1.1.   Just wonder rdkit validation itself can reduce how much gap?",
      "votes": null
    },
    {
      "id": "1337166",
      "postDate": "06/05/2021 13:04:27",
      "content": "<p>I think rotation and flip augmentation was useful to reduce LB/CV gap (also used for vertically long images in test inference). I also used multiple model voting but it does not help in my case maybe because for most records, predictions of models are all valid (same) or all invalid.</p>\n<p>I think rdkit validation for topK predictions by beam search is also useful but there was no left time to do that.</p>\n<p>When my score was higher than 1.0, rdkit validation improve LB score about 0.07 but in the later stages, the improvement was 0.02.</p>",
      "rawMarkdown": "I think rotation and flip augmentation was useful to reduce LB/CV gap (also used for vertically long images in test inference). I also used multiple model voting but it does not help in my case maybe because for most records, predictions of models are all valid (same) or all invalid.\n\nI think rdkit validation for topK predictions by beam search is also useful but there was no left time to do that.\n\nWhen my score was higher than 1.0, rdkit validation improve LB score about 0.07 but in the later stages, the improvement was 0.02.",
      "votes": null
    },
    {
      "id": "1337251",
      "postDate": "06/05/2021 13:59:48",
      "content": "<p>Thanks for your reply.</p>",
      "rawMarkdown": "Thanks for your reply.",
      "votes": null
    },
    {
      "id": "1337842",
      "postDate": "06/05/2021 23:55:24",
      "content": "<p>Congrats on the solo gold. Interesting that you found beam search or ensemble only useful on 20k images - did these have a distinctive characteristic? E.g. long InChI, extreme aspect ratio, larger/smaller than average? </p>",
      "rawMarkdown": "Congrats on the solo gold. Interesting that you found beam search or ensemble only useful on 20k images - did these have a distinctive characteristic? E.g. long InChI, extreme aspect ratio, larger/smaller than average?",
      "votes": null
    },
    {
      "id": "1337933",
      "postDate": "06/06/2021 03:43:20",
      "content": "<p>They are large molecules and have long InChI.</p>",
      "rawMarkdown": "They are large molecules and have long InChI.",
      "votes": null
    },
    {
      "id": "1340794",
      "postDate": "06/08/2021 09:11:48",
      "content": "<p>Congrats on becoming GM. you deserve it.</p>",
      "rawMarkdown": "Congrats on becoming GM. you deserve it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1336809,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/05/2021 08:17:34",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>! Delighted that you made it to GM with this one!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336903,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "06/05/2021 09:22:51",
      "content": "<p>Very nice! And congrats again for becoming GM. Well deserved:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336936,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/05/2021 09:52:27",
      "content": "<p>Thanks for the great write-up and congrats for the solo gold and becoming GM! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336951,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "06/05/2021 10:02:09",
      "content": "<p>Congratulation for becoming GM. I have big respect for solo teams in top 20.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336976,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/05/2021 10:32:19",
      "content": "<p>\"On the last day of the competition, I implemented beam search by referring to fairseq code.\"</p>\n<p>Congratulations on your good work!<br>\ncan you open-source k-beam search code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337072,
          "author_name": "charmq",
          "author_url": "",
          "post_date": "06/05/2021 11:55:44",
          "content": "<p>Thank you for your comment!<br>\nMy code is too dirty to share, but the beam search implementation is mostly a copy of fairseq repo and I think what you should implement yourself is reordering incremental states part. It is very easy to implement reordering like this.</p>\n<pre><code>def reorder_incremental_state(self, incremental_state, new_order):\n   for layer in self.text_decode.layer:\n      layer.self_attn.reorder_incremental_state(incremental_state, new_order)\n      layer.encoder_attn.reorder_incremental_state(incremental_state, new_order)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337016,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "06/05/2021 10:58:15",
      "content": "<p>Congrats! You deserve it. Get some rest and feel better :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337100,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "06/05/2021 12:16:31",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for solo gold and become GM! Thanks for sharing solution ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337117,
      "author_name": "akirasosa",
      "author_url": "",
      "post_date": "06/05/2021 12:31:40",
      "content": "<p>Congrats GM!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337152,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/05/2021 12:52:40",
      "content": "<p>Thanks for sharing.<br>\nbeside rdkit validation postprocess, what other process did you use to reduce LB/CV gap?  most other toppers use both rdkit validation and  multiple model voting etc. to correct invalid prediction.</p>\n<p>My LB/CV are 2.35/1.1.   Just wonder rdkit validation itself can reduce how much gap?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337166,
          "author_name": "charmq",
          "author_url": "",
          "post_date": "06/05/2021 13:04:27",
          "content": "<p>I think rotation and flip augmentation was useful to reduce LB/CV gap (also used for vertically long images in test inference). I also used multiple model voting but it does not help in my case maybe because for most records, predictions of models are all valid (same) or all invalid.</p>\n<p>I think rdkit validation for topK predictions by beam search is also useful but there was no left time to do that.</p>\n<p>When my score was higher than 1.0, rdkit validation improve LB score about 0.07 but in the later stages, the improvement was 0.02.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337251,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "06/05/2021 13:59:48",
          "content": "<p>Thanks for your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337842,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "06/05/2021 23:55:24",
      "content": "<p>Congrats on the solo gold. Interesting that you found beam search or ensemble only useful on 20k images - did these have a distinctive characteristic? E.g. long InChI, extreme aspect ratio, larger/smaller than average? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1337933,
          "author_name": "charmq",
          "author_url": "",
          "post_date": "06/06/2021 03:43:20",
          "content": "<p>They are large molecules and have long InChI.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1340794,
      "author_name": "sohailds",
      "author_url": "",
      "post_date": "06/08/2021 09:11:48",
      "content": "<p>Congrats on becoming GM. you deserve it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1336729": "Congratulations to all the winners and thank the competition hosts for holding such a wonderful competition.\nAnd especially thanks to @hengck23 for sharing the great ideas and code. I think his post made this competition very active.\n\nThis competition was very tough for a solo kaggler like me and I was very inspired by other solo kagglers ( @fergusoci, @AkiraSosa, @nofreewill) on LB. Finally, I got solo gold and was able to become a Grandmaster at the expense of health and many holidays!\n\n\n# Models\nI chose two models to train, SwinTransformer with image size 384x384 and patchwise TNT and I implemented it by referring to @hengck23's codes. \nI used hyper parameter embed_dim=768, in_dim=48 for TNT and this is larger than TNT-B.\n\nSwinTransformer and TNT both achieved LB~0.8 as single models and their CV were LD=0.80/0.80, CELoss=0.0030/0.0020\n\n\n# Data\nI generated about 10M images using all the InChIs in `extra_approved_InChIs.csv` using @stainsby's [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3). About 13M images are used in training. In the early stage of training, the loss for extra images was bad, but after long training, the loss for extra images and valid score improved. I trained SwinTransformer 15 epochs and TNT 20 epochs. SwinTransformer took 20 hours for 1 epoch and TNT took 12 hours for 1 epoch with 8 x V100.\n\n\n# Ensemble\nI ensembled 8 models in each decoder step (weighted average of predict_proba). Models are trained competition data+extra data or finetuned by only competition data after that or different fold. It seems important to set total weight of SwinTransformer and TNT equally. CV score of ensemble model was LD=0.60 and CELoss=0.0013 and LB:0.63\n\n\n# Beam Search\nOn the last day of the competition, I implemented beam search by referring to fairseq [code](https://github.com/pytorch/fairseq/blob/master/fairseq/sequence_generator.py). \nBy performing beam search with beam size 16, I could get LB:0.60.\nIn the later stages of the competition, I found that the prediction discrepancies of the different models of test data were only about 20k. Therefore, beam search or ensemble inference was performed only on 20k images.\n\n\n# PostProcess\nI always used the rdkit validation postprocess shared by @nofreewill.\nPost-processing improvement diminished with score improvement, but a little improvement remained even at the end.\n\n\n# What didn't work\n- I trained a generative model by using InChIs from the  competition training data and generated 8M images(molecules), but it did not help improve CV.\n- I trained the object detection model to detect atoms and bonds like the DACON solution and overlaid bboxes on the images, but it did not help improve CV.\n- TTA(rotation, flip)",
    "1336809": "Congratulations, @charmq! Delighted that you made it to GM with this one!",
    "1336903": "Very nice! And congrats again for becoming GM. Well deserved:)",
    "1336936": "Thanks for the great write-up and congrats for the solo gold and becoming GM!",
    "1336951": "Congratulation for becoming GM. I have big respect for solo teams in top 20.",
    "1336976": "\"On the last day of the competition, I implemented beam search by referring to fairseq code.\"\n\nCongratulations on your good work!\ncan you open-source k-beam search code?",
    "1337016": "Congrats! You deserve it. Get some rest and feel better :)",
    "1337072": "Thank you for your comment!\nMy code is too dirty to share, but the beam search implementation is mostly a copy of fairseq repo and I think what you should implement yourself is reordering incremental states part. It is very easy to implement reordering like this.\n\n~~~\ndef reorder_incremental_state(self, incremental_state, new_order):\n   for layer in self.text_decode.layer:\n      layer.self_attn.reorder_incremental_state(incremental_state, new_order)\n      layer.encoder_attn.reorder_incremental_state(incremental_state, new_order)\n~~~",
    "1337100": "Congrats @charmq for solo gold and become GM! Thanks for sharing solution ;)",
    "1337117": "Congrats GM!",
    "1337152": "Thanks for sharing.\nbeside rdkit validation postprocess, what other process did you use to reduce LB/CV gap?  most other toppers use both rdkit validation and  multiple model voting etc. to correct invalid prediction.\n\nMy LB/CV are 2.35/1.1.   Just wonder rdkit validation itself can reduce how much gap?",
    "1337166": "I think rotation and flip augmentation was useful to reduce LB/CV gap (also used for vertically long images in test inference). I also used multiple model voting but it does not help in my case maybe because for most records, predictions of models are all valid (same) or all invalid.\n\nI think rdkit validation for topK predictions by beam search is also useful but there was no left time to do that.\n\nWhen my score was higher than 1.0, rdkit validation improve LB score about 0.07 but in the later stages, the improvement was 0.02.",
    "1337251": "Thanks for your reply.",
    "1337842": "Congrats on the solo gold. Interesting that you found beam search or ensemble only useful on 20k images - did these have a distinctive characteristic? E.g. long InChI, extreme aspect ratio, larger/smaller than average?",
    "1337933": "They are large molecules and have long InChI.",
    "1340794": "Congrats on becoming GM. you deserve it."
  },
  "source": "meta"
}