{
  "id": 243779,
  "title": "7th Place Solution : My First Gold And I am in Tears",
  "url": "/competitions/bms-molecular-translation/writeups/hungry-for-gold-7th-place-solution-my-first-gold-a",
  "author_name": "",
  "post_date": "2021-06-04T07:25:54.160Z",
  "votes": 102,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>We would like to thank kaggle and the organizers for such a lovely competition. My learning in this has been immense , while a lot of people might criticize this competition to be a  hardware struggle , for me it more of a learning experience than a GPU pain mainly due to my lovely teammates <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/nishcaydnk\" target=\"_blank\">@nishcaydnk</a> who allowed and encouraged me to explore more and more.</p>\n<p>I always thought that I understood transformers completely but all that was shaken in this competition. We would like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for being so humble in sharing the transformer to transformer baseline with us , without that this medal would not have been possible . The code was easy to understand and a lot of our concepts got cleared during this competition . There are a lot of things we were still not able to implement like Beam search for our transformer model , but at the end of the day we are very happy with the result.</p>\n<p>We ( <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , me and <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a>) started with the CNN encoder with LSTM decoder like everyone else , quickly made our way to GRU , then we were joined by Rajneesh who brought in the idea of transformer to transformer . We fought hard to get to 1.30 then we were joined <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> and since then it has been an a rolercoaster ride . </p>\n<p>The last day was completely intense for us , we were at 0.65 and we knew we need something other than rdkit post-processing to get into gold , the only thing left was to do logits ensemble at every time step , we were unsuccesful at that before but nevertheless we tried it and it worked , our ensemble gave (0.66 lb 0.71cv) submitted at <b>11:30 UTC </b> and then after rdkit post-process we went to (0.60lb cv 0.66) submitted at <b> 11:45 UTC</b> and jumped to 7th place where we ended as well.</p>\n<p>Our final solution consists four major parts :</p>\n<ul>\n<li>Training transformer to transformer effectively </li>\n<li>Logit Level Ensemble at each time step</li>\n<li>Rdkit Post Processing</li>\n<li>Most Important : Analysis of all submissions and validation INCHI's</li>\n</ul>\n<h2>Models Used :</h2>\n<p>Our final ensemble consisted of </p>\n<ul>\n<li>Encoder - 2 ViT 384 Base, Decoder - Vanilla transformer  ( Cv-0.85 , Lb -0.76) and (Cv-0.86 , lb-0.77) respectively</li>\n<li>Encoder - 1 Swin Transformer 384 , Decoder -  Vanilla transformer  (Cv - 0.99 , LB -0.96)</li>\n</ul>\n<h2>Model Training :</h2>\n<p>We started ViT Training similar to heng's using lookahead optimizer and manually decreasing LR's . We made use of increasing resolution training and trained first with 224 Img_size , then fine-tuned the same model to 384 then finally to 448. This gave us a good amount of boost and we were able to reach 1.27 lb but it still wasn't enough . After analyzing the valid predictions we realized that the models kept failing on very long chained molecules and molecules with noises , our hypothesis was it was overfitting to the normal ones . Owing to the noisy images and our findings we decided to add Label Smoothing and Cutout to training and it worked , we were able to achieve (lb 0.77 ) with the same ViT model . After that we tried with 448 Image size with the same strategy however it didn't work so well and hence was excluded from final ensemble . We trained another ViT model with different LR and increased Random Scale as augmentation and were able to scoer (lb 0.76)</p>\n<p>We then started the search for diversifying our model zoo and moved on train SWIN and transformer-in-transformer with the same training Strategy , however Swin Beat tnt and scored (0.96 lb) as compared to tnt (1.07 lb)</p>\n<h2>RDkit Post-Processing</h2>\n<p>Everybody was normalizing the predicitons and we were no different , but only after reading <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> post we dug deeper and realized what we were missing.</p>\n<blockquote>\n  <p>def normalize_inchi(inchi):<br>\n      try:<br>\n          mol = Chem.MolFromInchi(inchi)<br>\n      except:<br>\n          pass<br>\n      if mol is None:<br>\n          return inchi,'invalid'<br>\n      else:<br>\n          try: <br>\n              norm = Chem.MolToInchi(mol)<br>\n              if norm == inchi:<br>\n                  return norm,'valid'<br>\n              else:<br>\n                  return norm,'modified'<br>\n          except: return inchi,'segmentation_failure'</p>\n</blockquote>\n<p>After adding this code to the normalization script we were able to get the status as well . We tested out the hypothesis that the places where INCHI were suggested invalid by rdkit were really wrong hence to replace them with valid predictions from different other model's prediction were a sure shot to improving score . We didn't stop here and dug even deeper to realize that the invalid INCHI's were coming from where the INCHI length was very long . Hence we fine-tuned our best tnt model on images having INCHi length greater than 150 and it gave great cv (0.99 previously 1.09) however due to lack of time we didn';t test lb and decided to add in post-processing to replace invalid INCHI's only</p>\n<h2>Models used in Post-Process to replace Invalid INCHI's</h2>\n<ul>\n<li>Tnt 336 Model </li>\n<li>Vit 448 Model</li>\n<li>Tnt 448 Model</li>\n<li>prunedeffnetb3 + GRU Model</li>\n<li>Effnetb4 + LSTM Model</li>\n<li>Effnetb7 + GRU Model</li>\n</ul>\n<h2>Logits Ensemble</h2>\n<p>I started coding Logit Ensemble on the last day and by the time it was ready only 9 hours were left in the competition , it gave a cv of 0.71 and hence we gave it our best to make it complete before 2 hours of deadline . Things fell into place and we were able to complete it in time.</p>\n<p>We realized that in heng's code the final logit gave an output of <b>(bs,vocab_size) , so untill the max_length and the vocab size of models are same</b> we can at every time step ensemble the predictions , and since we didn't have a lot of time , we hard-coded it , something like this :</p>\n<blockquote>\n  <p>class EnsembleNet(nn.Module):</p>\n</blockquote>\n<pre><code>def __init__(self, model_1, model_2, model_3):\n    super(EnsembleNet, self).__init__()\n    self.model_1 = model_1\n    self.model_2 = model_2\n    self.model_3 = model_3\n\n    self.model_1.load_state_dict(torch.load(initial_checkpoint1)['state_dict'], strict=True)\n    self.model_2.load_state_dict(torch.load(initial_checkpoint2)['state_dict'], strict=True)\n    self.model_3.load_state_dict(torch.load(initial_checkpoint3)['state_dict'], strict=True)\n\n    #self.model_1 = torch.jit.script(self.model_1)\n    #self.model_2 = torch.jit.script(self.model_2)\n\ndef forward(self, image):\n\n    image_dim = 768\n    text_dim = 768\n    decoder_dim = 768\n    num_layer = 4 #3\n    num_head = 8\n    ff_dim = 2048#1024\n\n    STOI = {\n        '&lt;sos&gt;': 190,\n        '&lt;eos&gt;': 191,\n        '&lt;pad&gt;': 192,\n        #'&lt;mask&gt;': 193,\n    }\n\n    image_size = 384\n    vocab_size = 193#194\n    max_length = 280 #300  # 275\n\n    device = image.device\n    batch_size = len(image)\n\n    image_embed_1 = self.model_1.cnn(image) #(bs,img_len,image_dim)\n    image_embed_1 = self.model_1.image_encode(image_embed_1).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n    image_embed_2 = self.model_2.cnn(image) #(bs,img_len,image_dim)\n    image_embed_2 = self.model_2.image_encode(image_embed_2).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n    image_embed_3 = self.model_3.cnn(image) #(bs,img_len,image_dim)\n    image_embed_3 = self.model_3.image_encode(image_embed_3).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n\n    token = torch.full((batch_size, max_length), STOI['&lt;pad&gt;'], dtype=torch.long, device=device) # (batch_size,max_len) \n    text_pos_1 = self.model_1.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    text_pos_2 = self.model_2.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    text_pos_3 = self.model_3.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    token[:, 0] = STOI['&lt;sos&gt;']\n    # -------------------------------------\n    eos = STOI['&lt;eos&gt;']\n    pad = STOI['&lt;pad&gt;']\n    # fast version\n    if 1:\n        # incremental_state = {}\n        incremental_state1 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        incremental_state2 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        incremental_state3 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        for t in range(max_length - 1):\n            last_token_1 = token[:, t] # take the whole batch's t'th token\n            text_embed_1 = self.model_1.token_embed(last_token_1) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_1 = text_embed_1 + text_pos_1[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_1 = text_embed_1.reshape(1, batch_size, 768)\n\n            last_token_2 = token[:, t] # take the whole batch's t'th token\n            text_embed_2 = self.model_2.token_embed(last_token_2) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_2 = text_embed_2 + text_pos_2[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_2 = text_embed_2.reshape(1, batch_size, 768)\n\n            last_token_3 = token[:, t] # take the whole batch's t'th token\n            text_embed_3 = self.model_3.token_embed(last_token_3) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_3 = text_embed_3 + text_pos_3[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_3 = text_embed_3.reshape(1, batch_size, 1024)\n\n            #text_embed ---&gt; 1,bs,text_dim(768)\n            #image_embed ---&gt; img_pos_embed,bs,image_im\n            x_1 = self.model_1.text_decode.forward_one(text_embed_1, image_embed_1, incremental_state1)\n            x_2 = self.model_2.text_decode.forward_one(text_embed_2, image_embed_2, incremental_state2)\n            x_3 = self.model_3.text_decode.forward_one(text_embed_3, image_embed_3, incremental_state3)\n            ## x -----&gt; (1,bs,text_dim)\n            x_1 = x_1.reshape(batch_size, 768)\n            x_2 = x_2.reshape(batch_size, 768)\n            x_3 = x_3.reshape(batch_size, 1024)\n            ## x -----&gt; (bs,decoder_dim)\n            l_1 = self.model_1.logit(x_1)\n            l_2 = self.model_2.logit(x_2)\n            l_3 = self.model_3.logit(x_3)\n            l = (l_1 + l_2 + l_3)/3\n            ## l -----&gt; (bs,num_classes)\n            k = torch.argmax(l, -1)\n            token[:, t + 1] = k\n            if ((k == eos) | (k == pad)).all():\n                break\n\n    predict = token[:, 1:]\n    return predict\n</code></pre>\n<h2>Final Hard-Voting</h2>\n<p>One Last Layer that we also added was hard-voting using all our submission (11 subs) . After doing RDkit Post Processing , we found that approx 11k invalid still remained and since we were not able to implement beam search for those invalid images we applied hard voting using all our model's submission ( <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> 's Idea)</p>\n<h1>Gratitude</h1>\n<p>This competition will remain closest to my heart , we have stayed up all night , its 6 AM in India right now , I will never forget how we passed the time from 1 to 4 while the ensemble inference was running . Congrats to <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> for becoming a master and  <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> for their second gold . It was a collective team effort and I couldn;t thank you all enough for taking me in your team . </p>\n<p>Its late , I will clean up the code soon and post here  very soon . To be continued ….<br>\nI hope you liked our solution . Thanks for reading </p>",
  "messages": [
    {
      "id": "1335017",
      "postDate": "06/04/2021 00:51:05",
      "content": "<p>Hi all, </p>\n<p>We would like to thank kaggle and the organizers for such a lovely competition. My learning in this has been immense , while a lot of people might criticize this competition to be a  hardware struggle , for me it more of a learning experience than a GPU pain mainly due to my lovely teammates <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/nishcaydnk\" target=\"_blank\">@nishcaydnk</a> who allowed and encouraged me to explore more and more.</p>\n<p>I always thought that I understood transformers completely but all that was shaken in this competition. We would like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for being so humble in sharing the transformer to transformer baseline with us , without that this medal would not have been possible . The code was easy to understand and a lot of our concepts got cleared during this competition . There are a lot of things we were still not able to implement like Beam search for our transformer model , but at the end of the day we are very happy with the result.</p>\n<p>We ( <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , me and <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a>) started with the CNN encoder with LSTM decoder like everyone else , quickly made our way to GRU , then we were joined by Rajneesh who brought in the idea of transformer to transformer . We fought hard to get to 1.30 then we were joined <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> and since then it has been an a rolercoaster ride . </p>\n<p>The last day was completely intense for us , we were at 0.65 and we knew we need something other than rdkit post-processing to get into gold , the only thing left was to do logits ensemble at every time step , we were unsuccesful at that before but nevertheless we tried it and it worked , our ensemble gave (0.66 lb 0.71cv) submitted at <b>11:30 UTC </b> and then after rdkit post-process we went to (0.60lb cv 0.66) submitted at <b> 11:45 UTC</b> and jumped to 7th place where we ended as well.</p>\n<p>Our final solution consists four major parts :</p>\n<ul>\n<li>Training transformer to transformer effectively </li>\n<li>Logit Level Ensemble at each time step</li>\n<li>Rdkit Post Processing</li>\n<li>Most Important : Analysis of all submissions and validation INCHI's</li>\n</ul>\n<h2>Models Used :</h2>\n<p>Our final ensemble consisted of </p>\n<ul>\n<li>Encoder - 2 ViT 384 Base, Decoder - Vanilla transformer  ( Cv-0.85 , Lb -0.76) and (Cv-0.86 , lb-0.77) respectively</li>\n<li>Encoder - 1 Swin Transformer 384 , Decoder -  Vanilla transformer  (Cv - 0.99 , LB -0.96)</li>\n</ul>\n<h2>Model Training :</h2>\n<p>We started ViT Training similar to heng's using lookahead optimizer and manually decreasing LR's . We made use of increasing resolution training and trained first with 224 Img_size , then fine-tuned the same model to 384 then finally to 448. This gave us a good amount of boost and we were able to reach 1.27 lb but it still wasn't enough . After analyzing the valid predictions we realized that the models kept failing on very long chained molecules and molecules with noises , our hypothesis was it was overfitting to the normal ones . Owing to the noisy images and our findings we decided to add Label Smoothing and Cutout to training and it worked , we were able to achieve (lb 0.77 ) with the same ViT model . After that we tried with 448 Image size with the same strategy however it didn't work so well and hence was excluded from final ensemble . We trained another ViT model with different LR and increased Random Scale as augmentation and were able to scoer (lb 0.76)</p>\n<p>We then started the search for diversifying our model zoo and moved on train SWIN and transformer-in-transformer with the same training Strategy , however Swin Beat tnt and scored (0.96 lb) as compared to tnt (1.07 lb)</p>\n<h2>RDkit Post-Processing</h2>\n<p>Everybody was normalizing the predicitons and we were no different , but only after reading <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> post we dug deeper and realized what we were missing.</p>\n<blockquote>\n  <p>def normalize_inchi(inchi):<br>\n      try:<br>\n          mol = Chem.MolFromInchi(inchi)<br>\n      except:<br>\n          pass<br>\n      if mol is None:<br>\n          return inchi,'invalid'<br>\n      else:<br>\n          try: <br>\n              norm = Chem.MolToInchi(mol)<br>\n              if norm == inchi:<br>\n                  return norm,'valid'<br>\n              else:<br>\n                  return norm,'modified'<br>\n          except: return inchi,'segmentation_failure'</p>\n</blockquote>\n<p>After adding this code to the normalization script we were able to get the status as well . We tested out the hypothesis that the places where INCHI were suggested invalid by rdkit were really wrong hence to replace them with valid predictions from different other model's prediction were a sure shot to improving score . We didn't stop here and dug even deeper to realize that the invalid INCHI's were coming from where the INCHI length was very long . Hence we fine-tuned our best tnt model on images having INCHi length greater than 150 and it gave great cv (0.99 previously 1.09) however due to lack of time we didn';t test lb and decided to add in post-processing to replace invalid INCHI's only</p>\n<h2>Models used in Post-Process to replace Invalid INCHI's</h2>\n<ul>\n<li>Tnt 336 Model </li>\n<li>Vit 448 Model</li>\n<li>Tnt 448 Model</li>\n<li>prunedeffnetb3 + GRU Model</li>\n<li>Effnetb4 + LSTM Model</li>\n<li>Effnetb7 + GRU Model</li>\n</ul>\n<h2>Logits Ensemble</h2>\n<p>I started coding Logit Ensemble on the last day and by the time it was ready only 9 hours were left in the competition , it gave a cv of 0.71 and hence we gave it our best to make it complete before 2 hours of deadline . Things fell into place and we were able to complete it in time.</p>\n<p>We realized that in heng's code the final logit gave an output of <b>(bs,vocab_size) , so untill the max_length and the vocab size of models are same</b> we can at every time step ensemble the predictions , and since we didn't have a lot of time , we hard-coded it , something like this :</p>\n<blockquote>\n  <p>class EnsembleNet(nn.Module):</p>\n</blockquote>\n<pre><code>def __init__(self, model_1, model_2, model_3):\n    super(EnsembleNet, self).__init__()\n    self.model_1 = model_1\n    self.model_2 = model_2\n    self.model_3 = model_3\n\n    self.model_1.load_state_dict(torch.load(initial_checkpoint1)['state_dict'], strict=True)\n    self.model_2.load_state_dict(torch.load(initial_checkpoint2)['state_dict'], strict=True)\n    self.model_3.load_state_dict(torch.load(initial_checkpoint3)['state_dict'], strict=True)\n\n    #self.model_1 = torch.jit.script(self.model_1)\n    #self.model_2 = torch.jit.script(self.model_2)\n\ndef forward(self, image):\n\n    image_dim = 768\n    text_dim = 768\n    decoder_dim = 768\n    num_layer = 4 #3\n    num_head = 8\n    ff_dim = 2048#1024\n\n    STOI = {\n        '&lt;sos&gt;': 190,\n        '&lt;eos&gt;': 191,\n        '&lt;pad&gt;': 192,\n        #'&lt;mask&gt;': 193,\n    }\n\n    image_size = 384\n    vocab_size = 193#194\n    max_length = 280 #300  # 275\n\n    device = image.device\n    batch_size = len(image)\n\n    image_embed_1 = self.model_1.cnn(image) #(bs,img_len,image_dim)\n    image_embed_1 = self.model_1.image_encode(image_embed_1).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n    image_embed_2 = self.model_2.cnn(image) #(bs,img_len,image_dim)\n    image_embed_2 = self.model_2.image_encode(image_embed_2).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n    image_embed_3 = self.model_3.cnn(image) #(bs,img_len,image_dim)\n    image_embed_3 = self.model_3.image_encode(image_embed_3).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n\n    token = torch.full((batch_size, max_length), STOI['&lt;pad&gt;'], dtype=torch.long, device=device) # (batch_size,max_len) \n    text_pos_1 = self.model_1.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    text_pos_2 = self.model_2.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    text_pos_3 = self.model_3.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n    token[:, 0] = STOI['&lt;sos&gt;']\n    # -------------------------------------\n    eos = STOI['&lt;eos&gt;']\n    pad = STOI['&lt;pad&gt;']\n    # fast version\n    if 1:\n        # incremental_state = {}\n        incremental_state1 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        incremental_state2 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        incremental_state3 = torch.jit.annotate(\n            Dict[str, Dict[str, Optional[torch.Tensor]]],\n            torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n        )\n        for t in range(max_length - 1):\n            last_token_1 = token[:, t] # take the whole batch's t'th token\n            text_embed_1 = self.model_1.token_embed(last_token_1) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_1 = text_embed_1 + text_pos_1[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_1 = text_embed_1.reshape(1, batch_size, 768)\n\n            last_token_2 = token[:, t] # take the whole batch's t'th token\n            text_embed_2 = self.model_2.token_embed(last_token_2) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_2 = text_embed_2 + text_pos_2[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_2 = text_embed_2.reshape(1, batch_size, 768)\n\n            last_token_3 = token[:, t] # take the whole batch's t'th token\n            text_embed_3 = self.model_3.token_embed(last_token_3) #[bs,text_dim] Generate embedding for the t'th token\n            text_embed_3 = text_embed_3 + text_pos_3[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n            text_embed_3 = text_embed_3.reshape(1, batch_size, 1024)\n\n            #text_embed ---&gt; 1,bs,text_dim(768)\n            #image_embed ---&gt; img_pos_embed,bs,image_im\n            x_1 = self.model_1.text_decode.forward_one(text_embed_1, image_embed_1, incremental_state1)\n            x_2 = self.model_2.text_decode.forward_one(text_embed_2, image_embed_2, incremental_state2)\n            x_3 = self.model_3.text_decode.forward_one(text_embed_3, image_embed_3, incremental_state3)\n            ## x -----&gt; (1,bs,text_dim)\n            x_1 = x_1.reshape(batch_size, 768)\n            x_2 = x_2.reshape(batch_size, 768)\n            x_3 = x_3.reshape(batch_size, 1024)\n            ## x -----&gt; (bs,decoder_dim)\n            l_1 = self.model_1.logit(x_1)\n            l_2 = self.model_2.logit(x_2)\n            l_3 = self.model_3.logit(x_3)\n            l = (l_1 + l_2 + l_3)/3\n            ## l -----&gt; (bs,num_classes)\n            k = torch.argmax(l, -1)\n            token[:, t + 1] = k\n            if ((k == eos) | (k == pad)).all():\n                break\n\n    predict = token[:, 1:]\n    return predict\n</code></pre>\n<h2>Final Hard-Voting</h2>\n<p>One Last Layer that we also added was hard-voting using all our submission (11 subs) . After doing RDkit Post Processing , we found that approx 11k invalid still remained and since we were not able to implement beam search for those invalid images we applied hard voting using all our model's submission ( <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> 's Idea)</p>\n<h1>Gratitude</h1>\n<p>This competition will remain closest to my heart , we have stayed up all night , its 6 AM in India right now , I will never forget how we passed the time from 1 to 4 while the ensemble inference was running . Congrats to <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> for becoming a master and  <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> for their second gold . It was a collective team effort and I couldn;t thank you all enough for taking me in your team . </p>\n<p>Its late , I will clean up the code soon and post here  very soon . To be continued ….<br>\nI hope you liked our solution . Thanks for reading </p>",
      "rawMarkdown": "Hi all, \n\nWe would like to thank kaggle and the organizers for such a lovely competition. My learning in this has been immense , while a lot of people might criticize this competition to be a  hardware struggle , for me it more of a learning experience than a GPU pain mainly due to my lovely teammates @pheadrus @shivamcyborg @atsunorifujita @nishcaydnk who allowed and encouraged me to explore more and more.\n \nI always thought that I understood transformers completely but all that was shaken in this competition. We would like to thank @hengck23 for being so humble in sharing the transformer to transformer baseline with us , without that this medal would not have been possible . The code was easy to understand and a lot of our concepts got cleared during this competition . There are a lot of things we were still not able to implement like Beam search for our transformer model , but at the end of the day we are very happy with the result.\n\nWe ( @nischaydnk , me and @shivamcyborg) started with the CNN encoder with LSTM decoder like everyone else , quickly made our way to GRU , then we were joined by Rajneesh who brought in the idea of transformer to transformer . We fought hard to get to 1.30 then we were joined @atsunorifujita and since then it has been an a rolercoaster ride . \n\nThe last day was completely intense for us , we were at 0.65 and we knew we need something other than rdkit post-processing to get into gold , the only thing left was to do logits ensemble at every time step , we were unsuccesful at that before but nevertheless we tried it and it worked , our ensemble gave (0.66 lb 0.71cv) submitted at <b>11:30 UTC </b> and then after rdkit post-process we went to (0.60lb cv 0.66) submitted at <b> 11:45 UTC</b> and jumped to 7th place where we ended as well.\n\nOur final solution consists four major parts :\n* Training transformer to transformer effectively \n* Logit Level Ensemble at each time step\n* Rdkit Post Processing\n* Most Important : Analysis of all submissions and validation INCHI's\n\n## Models Used :\n\nOur final ensemble consisted of \n* Encoder - 2 ViT 384 Base, Decoder - Vanilla transformer  ( Cv-0.85 , Lb -0.76) and (Cv-0.86 , lb-0.77) respectively\n* Encoder - 1 Swin Transformer 384 , Decoder -  Vanilla transformer  (Cv - 0.99 , LB -0.96)\n\n## Model Training :\n\nWe started ViT Training similar to heng's using lookahead optimizer and manually decreasing LR's . We made use of increasing resolution training and trained first with 224 Img_size , then fine-tuned the same model to 384 then finally to 448. This gave us a good amount of boost and we were able to reach 1.27 lb but it still wasn't enough . After analyzing the valid predictions we realized that the models kept failing on very long chained molecules and molecules with noises , our hypothesis was it was overfitting to the normal ones . Owing to the noisy images and our findings we decided to add Label Smoothing and Cutout to training and it worked , we were able to achieve (lb 0.77 ) with the same ViT model . After that we tried with 448 Image size with the same strategy however it didn't work so well and hence was excluded from final ensemble . We trained another ViT model with different LR and increased Random Scale as augmentation and were able to scoer (lb 0.76)\n\nWe then started the search for diversifying our model zoo and moved on train SWIN and transformer-in-transformer with the same training Strategy , however Swin Beat tnt and scored (0.96 lb) as compared to tnt (1.07 lb)\n\n## RDkit Post-Processing\n\nEverybody was normalizing the predicitons and we were no different , but only after reading @nofreewill post we dug deeper and realized what we were missing.\n\n> def normalize_inchi(inchi):\n    try:\n        mol = Chem.MolFromInchi(inchi)\n    except:\n        pass\n    if mol is None:\n        return inchi,'invalid'\n    else:\n        try: \n            norm = Chem.MolToInchi(mol)\n            if norm == inchi:\n                return norm,'valid'\n            else:\n                return norm,'modified'\n        except: return inchi,'segmentation_failure'\n\nAfter adding this code to the normalization script we were able to get the status as well . We tested out the hypothesis that the places where INCHI were suggested invalid by rdkit were really wrong hence to replace them with valid predictions from different other model's prediction were a sure shot to improving score . We didn't stop here and dug even deeper to realize that the invalid INCHI's were coming from where the INCHI length was very long . Hence we fine-tuned our best tnt model on images having INCHi length greater than 150 and it gave great cv (0.99 previously 1.09) however due to lack of time we didn';t test lb and decided to add in post-processing to replace invalid INCHI's only\n\n## Models used in Post-Process to replace Invalid INCHI's\n\n* Tnt 336 Model \n* Vit 448 Model\n* Tnt 448 Model\n* prunedeffnetb3 + GRU Model\n* Effnetb4 + LSTM Model\n* Effnetb7 + GRU Model\n\n## Logits Ensemble\n\nI started coding Logit Ensemble on the last day and by the time it was ready only 9 hours were left in the competition , it gave a cv of 0.71 and hence we gave it our best to make it complete before 2 hours of deadline . Things fell into place and we were able to complete it in time.\n\nWe realized that in heng's code the final logit gave an output of <b>(bs,vocab_size) , so untill the max_length and the vocab size of models are same</b> we can at every time step ensemble the predictions , and since we didn't have a lot of time , we hard-coded it , something like this :\n\n> class EnsembleNet(nn.Module):\n    \n    def __init__(self, model_1, model_2, model_3):\n        super(EnsembleNet, self).__init__()\n        self.model_1 = model_1\n        self.model_2 = model_2\n        self.model_3 = model_3\n        \n        self.model_1.load_state_dict(torch.load(initial_checkpoint1)['state_dict'], strict=True)\n        self.model_2.load_state_dict(torch.load(initial_checkpoint2)['state_dict'], strict=True)\n        self.model_3.load_state_dict(torch.load(initial_checkpoint3)['state_dict'], strict=True)\n        \n        #self.model_1 = torch.jit.script(self.model_1)\n        #self.model_2 = torch.jit.script(self.model_2)\n        \n    def forward(self, image):\n        \n        image_dim = 768\n        text_dim = 768\n        decoder_dim = 768\n        num_layer = 4 #3\n        num_head = 8\n        ff_dim = 2048#1024\n\n        STOI = {\n            '<sos>': 190,\n            '<eos>': 191,\n            '<pad>': 192,\n            #'<mask>': 193,\n        }\n\n        image_size = 384\n        vocab_size = 193#194\n        max_length = 280 #300  # 275\n        \n        device = image.device\n        batch_size = len(image)\n        \n        image_embed_1 = self.model_1.cnn(image) #(bs,img_len,image_dim)\n        image_embed_1 = self.model_1.image_encode(image_embed_1).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        image_embed_2 = self.model_2.cnn(image) #(bs,img_len,image_dim)\n        image_embed_2 = self.model_2.image_encode(image_embed_2).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        image_embed_3 = self.model_3.cnn(image) #(bs,img_len,image_dim)\n        image_embed_3 = self.model_3.image_encode(image_embed_3).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        \n        token = torch.full((batch_size, max_length), STOI['<pad>'], dtype=torch.long, device=device) # (batch_size,max_len) \n        text_pos_1 = self.model_1.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        text_pos_2 = self.model_2.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        text_pos_3 = self.model_3.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        token[:, 0] = STOI['<sos>']\n        # -------------------------------------\n        eos = STOI['<eos>']\n        pad = STOI['<pad>']\n        # fast version\n        if 1:\n            # incremental_state = {}\n            incremental_state1 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            incremental_state2 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            incremental_state3 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            for t in range(max_length - 1):\n                last_token_1 = token[:, t] # take the whole batch's t'th token\n                text_embed_1 = self.model_1.token_embed(last_token_1) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_1 = text_embed_1 + text_pos_1[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_1 = text_embed_1.reshape(1, batch_size, 768)\n                \n                last_token_2 = token[:, t] # take the whole batch's t'th token\n                text_embed_2 = self.model_2.token_embed(last_token_2) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_2 = text_embed_2 + text_pos_2[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_2 = text_embed_2.reshape(1, batch_size, 768)\n                \n                last_token_3 = token[:, t] # take the whole batch's t'th token\n                text_embed_3 = self.model_3.token_embed(last_token_3) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_3 = text_embed_3 + text_pos_3[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_3 = text_embed_3.reshape(1, batch_size, 1024)\n                \n                #text_embed ---> 1,bs,text_dim(768)\n                #image_embed ---> img_pos_embed,bs,image_im\n                x_1 = self.model_1.text_decode.forward_one(text_embed_1, image_embed_1, incremental_state1)\n                x_2 = self.model_2.text_decode.forward_one(text_embed_2, image_embed_2, incremental_state2)\n                x_3 = self.model_3.text_decode.forward_one(text_embed_3, image_embed_3, incremental_state3)\n                ## x -----> (1,bs,text_dim)\n                x_1 = x_1.reshape(batch_size, 768)\n                x_2 = x_2.reshape(batch_size, 768)\n                x_3 = x_3.reshape(batch_size, 1024)\n                ## x -----> (bs,decoder_dim)\n                l_1 = self.model_1.logit(x_1)\n                l_2 = self.model_2.logit(x_2)\n                l_3 = self.model_3.logit(x_3)\n                l = (l_1 + l_2 + l_3)/3\n                ## l -----> (bs,num_classes)\n                k = torch.argmax(l, -1)\n                token[:, t + 1] = k\n                if ((k == eos) | (k == pad)).all():\n                    break\n        \n        predict = token[:, 1:]\n        return predict\n\n## Final Hard-Voting\n\nOne Last Layer that we also added was hard-voting using all our submission (11 subs) . After doing RDkit Post Processing , we found that approx 11k invalid still remained and since we were not able to implement beam search for those invalid images we applied hard voting using all our model's submission ( @nischaydnk 's Idea)\n\n# Gratitude\n\nThis competition will remain closest to my heart , we have stayed up all night , its 6 AM in India right now , I will never forget how we passed the time from 1 to 4 while the ensemble inference was running . Congrats to @pheadrus @shivamcyborg for becoming a master and  @atsunorifujita @nischaydnk for their second gold . It was a collective team effort and I couldn;t thank you all enough for taking me in your team . \n\nIts late , I will clean up the code soon and post here  very soon . To be continued ....\nI hope you liked our solution . Thanks for reading",
      "votes": null
    },
    {
      "id": "1335032",
      "postDate": "06/04/2021 01:13:52",
      "content": "<p>Congratulations on your gold medal <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> 🎉🎉</p>\n<p>The jump to the gold zone on the last day is amazing.😆</p>",
      "rawMarkdown": "Congratulations on your gold medal @nischaydnk  @shivamcyborg 🎉🎉\n\nThe jump to the gold zone on the last day is amazing.😆",
      "votes": null
    },
    {
      "id": "1335034",
      "postDate": "06/04/2021 01:17:04",
      "content": "<p>Congrats on you and your team's fist gold medal. Well done! :)<br>\n<a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "rawMarkdown": "Congrats on you and your team's fist gold medal. Well done! :)\n@tanulsingh077 @atsunorifujita @shivamcyborg @pheadrus @nischaydnk",
      "votes": null
    },
    {
      "id": "1335037",
      "postDate": "06/04/2021 01:24:52",
      "content": "<p>Congrats to competition master</p>",
      "rawMarkdown": "Congrats to competition master",
      "votes": null
    },
    {
      "id": "1335043",
      "postDate": "06/04/2021 01:32:41",
      "content": "<p>Congrats for the gold :)</p>",
      "rawMarkdown": "Congrats for the gold :)",
      "votes": null
    },
    {
      "id": "1335064",
      "postDate": "06/04/2021 01:59:18",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> . You did it bro, now competition master 🙌🙌🙌</p>",
      "rawMarkdown": "Congrats @tanulsingh077 . You did it bro, now competition master 🙌🙌🙌",
      "votes": null
    },
    {
      "id": "1335077",
      "postDate": "06/04/2021 02:31:13",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> on the gold and becoming a competition master !! Well deserved!!</p>",
      "rawMarkdown": "Congrats @tanulsingh077 on the gold and becoming a competition master !! Well deserved!!",
      "votes": null
    },
    {
      "id": "1335093",
      "postDate": "06/04/2021 02:47:33",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a></p>",
      "rawMarkdown": "Congrats @tanulsingh077",
      "votes": null
    },
    {
      "id": "1335117",
      "postDate": "06/04/2021 03:00:56",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a></p>",
      "rawMarkdown": "Congrats @tanulsingh077",
      "votes": null
    },
    {
      "id": "1335226",
      "postDate": "06/04/2021 05:16:06",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> for gold and becoming competition masters, an Amazing feat! Looking forward to learning from you guys 😄</p>",
      "rawMarkdown": "Congratulations @tanulsingh077 @nischaydnk @pheadrus @shivamcyborg @atsunorifujita for gold and becoming competition masters, an Amazing feat! Looking forward to learning from you guys 😄",
      "votes": null
    },
    {
      "id": "1335265",
      "postDate": "06/04/2021 05:48:47",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> , that was intense. 😄</p>",
      "rawMarkdown": "Thank you @dehokanta , that was intense. 😄",
      "votes": null
    },
    {
      "id": "1335356",
      "postDate": "06/04/2021 07:05:01",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/heroseo\" target=\"_blank\">@heroseo</a> , its time to move on to SETI I guess 😜</p>",
      "rawMarkdown": "Thanks a lot @heroseo , its time to move on to SETI I guess 😜",
      "votes": null
    },
    {
      "id": "1335358",
      "postDate": "06/04/2021 07:05:23",
      "content": "<p>Thank you </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "1335359",
      "postDate": "06/04/2021 07:05:35",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/Gao\" target=\"_blank\">@Gao</a></p>",
      "rawMarkdown": "Thanks a lot @Gao",
      "votes": null
    },
    {
      "id": "1335360",
      "postDate": "06/04/2021 07:05:54",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/kamal\" target=\"_blank\">@kamal</a> , we finally got our gold</p>",
      "rawMarkdown": "Thank you @kamal , we finally got our gold",
      "votes": null
    },
    {
      "id": "1335361",
      "postDate": "06/04/2021 07:06:22",
      "content": "<p>Thanks bro , it sure took long but finally it happened 😊</p>",
      "rawMarkdown": "Thanks bro , it sure took long but finally it happened 😊",
      "votes": null
    },
    {
      "id": "1335362",
      "postDate": "06/04/2021 07:06:43",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/atharva\" target=\"_blank\">@atharva</a></p>",
      "rawMarkdown": "Thank you @atharva",
      "votes": null
    },
    {
      "id": "1335425",
      "postDate": "06/04/2021 07:51:22",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> for becoming competition Master and to the whole team for the outstanding results 🎉🎉. Interesting to notice how you improved the training of transformers models by reducing overfitting with cutout and label smoothing. We should have definitely tried that too. looking forward to seeing your code to learn more about your approach</p>",
      "rawMarkdown": "Congratulation @tanulsingh077 for becoming competition Master and to the whole team for the outstanding results 🎉🎉. Interesting to notice how you improved the training of transformers models by reducing overfitting with cutout and label smoothing. We should have definitely tried that too. looking forward to seeing your code to learn more about your approach",
      "votes": null
    },
    {
      "id": "1335434",
      "postDate": "06/04/2021 08:01:13",
      "content": "<p>Congratz to you and your team ! It was about time the \"hungry for gold\" team got its meal :)</p>",
      "rawMarkdown": "Congratz to you and your team ! It was about time the \"hungry for gold\" team got its meal :)",
      "votes": null
    },
    {
      "id": "1335439",
      "postDate": "06/04/2021 08:08:39",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> and team! If I have to pick the thing that had the biggest impact it would be \"Label Smoothing and Cutout\" which got you from 1.27 to 0.77 LB. That's a HUGE boost. Would you agree? By the way what is this \"transformer to transformer\"? Are you referring to TNT?</p>\n<p>Again, congrats on the standings in this competition, and for levelling up!</p>",
      "rawMarkdown": "Great work @tanulsingh077 and team! If I have to pick the thing that had the biggest impact it would be \"Label Smoothing and Cutout\" which got you from 1.27 to 0.77 LB. That's a HUGE boost. Would you agree? By the way what is this \"transformer to transformer\"? Are you referring to TNT?\n\nAgain, congrats on the standings in this competition, and for levelling up!",
      "votes": null
    },
    {
      "id": "1335481",
      "postDate": "06/04/2021 08:43:19",
      "content": "<p>Congratz! Great job, you make it!</p>",
      "rawMarkdown": "Congratz! Great job, you make it!",
      "votes": null
    },
    {
      "id": "1335483",
      "postDate": "06/04/2021 08:44:09",
      "content": "<p>It always feels good when your idol appreciates you 😉 <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , thanks a lot <br>\nHungry for gold has a long way to go </p>",
      "rawMarkdown": "It always feels good when your idol appreciates you 😉 @theoviel , thanks a lot \nHungry for gold has a long way to go",
      "votes": null
    },
    {
      "id": "1335484",
      "postDate": "06/04/2021 08:46:28",
      "content": "<p>Yes , that really surprised us even , but even more crucial was the logits ensemble which increased our best score without PP from 0.77 ----&gt; 0.66 </p>\n<p>The thing that helped us the most was constantly monitoring and analyzing the validation INCHI and trying to figure out what and where the models were doing wrong</p>",
      "rawMarkdown": "Yes , that really surprised us even , but even more crucial was the logits ensemble which increased our best score without PP from 0.77 ----> 0.66 \n\nThe thing that helped us the most was constantly monitoring and analyzing the validation INCHI and trying to figure out what and where the models were doing wrong",
      "votes": null
    },
    {
      "id": "1335488",
      "postDate": "06/04/2021 08:47:39",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a> , label smoothing was fujita's idea and we were also surprised of how good it worked</p>",
      "rawMarkdown": "Thanks a lot @jonathanbesomi , label smoothing was fujita's idea and we were also surprised of how good it worked",
      "votes": null
    },
    {
      "id": "1335511",
      "postDate": "06/04/2021 09:07:50",
      "content": "<p>yes <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> by \"transformer to transformer\" we mean TNT and congratulations on your solo Silver medal 27th Place finish. 👍</p>",
      "rawMarkdown": "yes @alexandersoare by \"transformer to transformer\" we mean TNT and congratulations on your solo Silver medal 27th Place finish. 👍",
      "votes": null
    },
    {
      "id": "1335542",
      "postDate": "06/04/2021 09:21:29",
      "content": "<p>Do you train your models in local machine or cloud?</p>",
      "rawMarkdown": "Do you train your models in local machine or cloud?",
      "votes": null
    },
    {
      "id": "1335614",
      "postDate": "06/04/2021 10:12:55",
      "content": "<p>We train models in cloud. Some of us have gpu locally setup in PC to run experiments.</p>",
      "rawMarkdown": "We train models in cloud. Some of us have gpu locally setup in PC to run experiments.",
      "votes": null
    },
    {
      "id": "1335618",
      "postDate": "06/04/2021 10:22:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> !</p>",
      "rawMarkdown": "Thanks @shivamcyborg !",
      "votes": null
    },
    {
      "id": "1335669",
      "postDate": "06/04/2021 11:07:29",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and many many congratulations to you and your team on gold 🥇 medal 5th place finish.</p>",
      "rawMarkdown": "Thanks @haqishen and many many congratulations to you and your team on gold 🥇 medal 5th place finish.",
      "votes": null
    },
    {
      "id": "1335672",
      "postDate": "06/04/2021 11:11:07",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> , last day idea worked pretty well here on time. </p>",
      "rawMarkdown": "Thank you @dehokanta , last day idea worked pretty well here on time.",
      "votes": null
    },
    {
      "id": "1335681",
      "postDate": "06/04/2021 11:20:48",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/heroseo\" target=\"_blank\">@heroseo</a></p>",
      "rawMarkdown": "Thanks @heroseo",
      "votes": null
    },
    {
      "id": "1335696",
      "postDate": "06/04/2021 11:36:15",
      "content": "<p>Hi may I know which cloud provider you guys used for your model training? </p>",
      "rawMarkdown": "Hi may I know which cloud provider you guys used for your model training?",
      "votes": null
    },
    {
      "id": "1335730",
      "postDate": "06/04/2021 11:58:21",
      "content": "<p>It was GCP and google colab pro from time to time.</p>",
      "rawMarkdown": "It was GCP and google colab pro from time to time.",
      "votes": null
    },
    {
      "id": "1335767",
      "postDate": "06/04/2021 12:28:41",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>",
      "rawMarkdown": "Thanks a lot @haqishen",
      "votes": null
    },
    {
      "id": "1335883",
      "postDate": "06/04/2021 13:47:58",
      "content": "<p>Tears, immense joy. Kaggle can bring out the best in us 😉<br>\nCongrats on the last minute gold for you and team!</p>",
      "rawMarkdown": "Tears, immense joy. Kaggle can bring out the best in us 😉\nCongrats on the last minute gold for you and team!",
      "votes": null
    },
    {
      "id": "1336114",
      "postDate": "06/04/2021 16:21:56",
      "content": "<p>congrats for winning gold medal !!<br>\nCould you please explain what you meant by transformer to transformer..? Is it transformer on encoder and decoder part..?<br>\nif possible refer would be appreciated.</p>",
      "rawMarkdown": "congrats for winning gold medal !!\nCould you please explain what you meant by transformer to transformer..? Is it transformer on encoder and decoder part..?\nif possible refer would be appreciated.",
      "votes": null
    },
    {
      "id": "1336484",
      "postDate": "06/05/2021 00:58:20",
      "content": "<p>Congrats! Tears are a natural way of expressing the dedication and hardwork….More on your way</p>",
      "rawMarkdown": "Congrats! Tears are a natural way of expressing the dedication and hardwork....More on your way",
      "votes": null
    },
    {
      "id": "1336881",
      "postDate": "06/05/2021 09:06:13",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> </p>",
      "rawMarkdown": "Thanks @rafiko1",
      "votes": null
    },
    {
      "id": "1336890",
      "postDate": "06/05/2021 09:10:45",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/prashantkramadhari\" target=\"_blank\">@prashantkramadhari</a> transformer to transformer means transformer type encoder (like vision transformer , Swin or tnt ) and a transformer decoder (like vanilla transformer,BERT,etc)</p>",
      "rawMarkdown": "Yes @prashantkramadhari transformer to transformer means transformer type encoder (like vision transformer , Swin or tnt ) and a transformer decoder (like vanilla transformer,BERT,etc)",
      "votes": null
    },
    {
      "id": "1336900",
      "postDate": "06/05/2021 09:17:20",
      "content": "<p>Wow, Congrats Tanul!</p>",
      "rawMarkdown": "Wow, Congrats Tanul!",
      "votes": null
    },
    {
      "id": "1337904",
      "postDate": "06/06/2021 02:33:23",
      "content": "<p>Congratulations Mr_KnowNothing and team. Well done. I am happy that you got Gold. Great solution!</p>",
      "rawMarkdown": "Congratulations Mr_KnowNothing and team. Well done. I am happy that you got Gold. Great solution!",
      "votes": null
    },
    {
      "id": "1338100",
      "postDate": "06/06/2021 07:18:28",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , it feels great to be appreciated by you . I hope to continue this further </p>",
      "rawMarkdown": "Thanks a lot @cdeotte , it feels great to be appreciated by you . I hope to continue this further",
      "votes": null
    },
    {
      "id": "1338102",
      "postDate": "06/06/2021 07:19:09",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "1338151",
      "postDate": "06/06/2021 08:06:24",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> :)</p>",
      "rawMarkdown": "Thank you @cdeotte :)",
      "votes": null
    },
    {
      "id": "1338570",
      "postDate": "06/06/2021 14:51:23",
      "content": "<p>Congratulations, very well deserved!</p>",
      "rawMarkdown": "Congratulations, very well deserved!",
      "votes": null
    },
    {
      "id": "1534240",
      "postDate": "10/04/2021 18:21:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> for your achievements.  </p>",
      "rawMarkdown": "Congratulations @tanulsingh077 @nischaydnk @pheadrus @shivamcyborg @atsunorifujita for your achievements.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335032,
      "author_name": "dehokanta",
      "author_url": "",
      "post_date": "06/04/2021 01:13:52",
      "content": "<p>Congratulations on your gold medal <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> 🎉🎉</p>\n<p>The jump to the gold zone on the last day is amazing.😆</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335265,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "06/04/2021 05:48:47",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> , that was intense. 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335672,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 11:11:07",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> , last day idea worked pretty well here on time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335034,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "06/04/2021 01:17:04",
      "content": "<p>Congrats on you and your team's fist gold medal. Well done! :)<br>\n<a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1335356,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:05:01",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/heroseo\" target=\"_blank\">@heroseo</a> , its time to move on to SETI I guess 😜</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335681,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 11:20:48",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/heroseo\" target=\"_blank\">@heroseo</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335037,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "06/04/2021 01:24:52",
      "content": "<p>Congrats to competition master</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335043,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "06/04/2021 01:32:41",
      "content": "<p>Congrats for the gold :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335362,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:06:43",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/atharva\" target=\"_blank\">@atharva</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335064,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "06/04/2021 01:59:18",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> . You did it bro, now competition master 🙌🙌🙌</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335361,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:06:22",
          "content": "<p>Thanks bro , it sure took long but finally it happened 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335077,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "06/04/2021 02:31:13",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> on the gold and becoming a competition master !! Well deserved!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335360,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:05:54",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/kamal\" target=\"_blank\">@kamal</a> , we finally got our gold</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335093,
      "author_name": "sid00733",
      "author_url": "",
      "post_date": "06/04/2021 02:47:33",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335117,
      "author_name": "koukinn",
      "author_url": "",
      "post_date": "06/04/2021 03:00:56",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1335359,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:05:35",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/Gao\" target=\"_blank\">@Gao</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335226,
      "author_name": "lplenka",
      "author_url": "",
      "post_date": "06/04/2021 05:16:06",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> for gold and becoming competition masters, an Amazing feat! Looking forward to learning from you guys 😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335358,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 07:05:23",
          "content": "<p>Thank you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335425,
      "author_name": "jonathanbesomi",
      "author_url": "",
      "post_date": "06/04/2021 07:51:22",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> for becoming competition Master and to the whole team for the outstanding results 🎉🎉. Interesting to notice how you improved the training of transformers models by reducing overfitting with cutout and label smoothing. We should have definitely tried that too. looking forward to seeing your code to learn more about your approach</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335488,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 08:47:39",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a> , label smoothing was fujita's idea and we were also surprised of how good it worked</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335434,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "06/04/2021 08:01:13",
      "content": "<p>Congratz to you and your team ! It was about time the \"hungry for gold\" team got its meal :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335483,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 08:44:09",
          "content": "<p>It always feels good when your idol appreciates you 😉 <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , thanks a lot <br>\nHungry for gold has a long way to go </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335439,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "06/04/2021 08:08:39",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> and team! If I have to pick the thing that had the biggest impact it would be \"Label Smoothing and Cutout\" which got you from 1.27 to 0.77 LB. That's a HUGE boost. Would you agree? By the way what is this \"transformer to transformer\"? Are you referring to TNT?</p>\n<p>Again, congrats on the standings in this competition, and for levelling up!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335484,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 08:46:28",
          "content": "<p>Yes , that really surprised us even , but even more crucial was the logits ensemble which increased our best score without PP from 0.77 ----&gt; 0.66 </p>\n<p>The thing that helped us the most was constantly monitoring and analyzing the validation INCHI and trying to figure out what and where the models were doing wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335511,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 09:07:50",
          "content": "<p>yes <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> by \"transformer to transformer\" we mean TNT and congratulations on your solo Silver medal 27th Place finish. 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335618,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "06/04/2021 10:22:51",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335481,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "06/04/2021 08:43:19",
      "content": "<p>Congratz! Great job, you make it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335669,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 11:07:29",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and many many congratulations to you and your team on gold 🥇 medal 5th place finish.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335767,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/04/2021 12:28:41",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335542,
      "author_name": "adityatewary",
      "author_url": "",
      "post_date": "06/04/2021 09:21:29",
      "content": "<p>Do you train your models in local machine or cloud?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335614,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 10:12:55",
          "content": "<p>We train models in cloud. Some of us have gpu locally setup in PC to run experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335696,
          "author_name": "oatmeals",
          "author_url": "",
          "post_date": "06/04/2021 11:36:15",
          "content": "<p>Hi may I know which cloud provider you guys used for your model training? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335730,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/04/2021 11:58:21",
          "content": "<p>It was GCP and google colab pro from time to time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335883,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "06/04/2021 13:47:58",
      "content": "<p>Tears, immense joy. Kaggle can bring out the best in us 😉<br>\nCongrats on the last minute gold for you and team!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336881,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/05/2021 09:06:13",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336114,
      "author_name": "prashantkramadhari",
      "author_url": "",
      "post_date": "06/04/2021 16:21:56",
      "content": "<p>congrats for winning gold medal !!<br>\nCould you please explain what you meant by transformer to transformer..? Is it transformer on encoder and decoder part..?<br>\nif possible refer would be appreciated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336890,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/05/2021 09:10:45",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/prashantkramadhari\" target=\"_blank\">@prashantkramadhari</a> transformer to transformer means transformer type encoder (like vision transformer , Swin or tnt ) and a transformer decoder (like vanilla transformer,BERT,etc)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336484,
      "author_name": "drprakashmuthudoss",
      "author_url": "",
      "post_date": "06/05/2021 00:58:20",
      "content": "<p>Congrats! Tears are a natural way of expressing the dedication and hardwork….More on your way</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336900,
      "author_name": "ankitsajwan",
      "author_url": "",
      "post_date": "06/05/2021 09:17:20",
      "content": "<p>Wow, Congrats Tanul!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337904,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/06/2021 02:33:23",
      "content": "<p>Congratulations Mr_KnowNothing and team. Well done. I am happy that you got Gold. Great solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1338100,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/06/2021 07:18:28",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , it feels great to be appreciated by you . I hope to continue this further </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338151,
          "author_name": "shivamcyborg",
          "author_url": "",
          "post_date": "06/06/2021 08:06:24",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1338102,
      "author_name": "mdsumonhossain",
      "author_url": "",
      "post_date": "06/06/2021 07:19:09",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338570,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "06/06/2021 14:51:23",
      "content": "<p>Congratulations, very well deserved!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1534240,
      "author_name": "yvonnef",
      "author_url": "",
      "post_date": "10/04/2021 18:21:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> for your achievements.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335017": "Hi all, \n\nWe would like to thank kaggle and the organizers for such a lovely competition. My learning in this has been immense , while a lot of people might criticize this competition to be a  hardware struggle , for me it more of a learning experience than a GPU pain mainly due to my lovely teammates @pheadrus @shivamcyborg @atsunorifujita @nishcaydnk who allowed and encouraged me to explore more and more.\n \nI always thought that I understood transformers completely but all that was shaken in this competition. We would like to thank @hengck23 for being so humble in sharing the transformer to transformer baseline with us , without that this medal would not have been possible . The code was easy to understand and a lot of our concepts got cleared during this competition . There are a lot of things we were still not able to implement like Beam search for our transformer model , but at the end of the day we are very happy with the result.\n\nWe ( @nischaydnk , me and @shivamcyborg) started with the CNN encoder with LSTM decoder like everyone else , quickly made our way to GRU , then we were joined by Rajneesh who brought in the idea of transformer to transformer . We fought hard to get to 1.30 then we were joined @atsunorifujita and since then it has been an a rolercoaster ride . \n\nThe last day was completely intense for us , we were at 0.65 and we knew we need something other than rdkit post-processing to get into gold , the only thing left was to do logits ensemble at every time step , we were unsuccesful at that before but nevertheless we tried it and it worked , our ensemble gave (0.66 lb 0.71cv) submitted at <b>11:30 UTC </b> and then after rdkit post-process we went to (0.60lb cv 0.66) submitted at <b> 11:45 UTC</b> and jumped to 7th place where we ended as well.\n\nOur final solution consists four major parts :\n* Training transformer to transformer effectively \n* Logit Level Ensemble at each time step\n* Rdkit Post Processing\n* Most Important : Analysis of all submissions and validation INCHI's\n\n## Models Used :\n\nOur final ensemble consisted of \n* Encoder - 2 ViT 384 Base, Decoder - Vanilla transformer  ( Cv-0.85 , Lb -0.76) and (Cv-0.86 , lb-0.77) respectively\n* Encoder - 1 Swin Transformer 384 , Decoder -  Vanilla transformer  (Cv - 0.99 , LB -0.96)\n\n## Model Training :\n\nWe started ViT Training similar to heng's using lookahead optimizer and manually decreasing LR's . We made use of increasing resolution training and trained first with 224 Img_size , then fine-tuned the same model to 384 then finally to 448. This gave us a good amount of boost and we were able to reach 1.27 lb but it still wasn't enough . After analyzing the valid predictions we realized that the models kept failing on very long chained molecules and molecules with noises , our hypothesis was it was overfitting to the normal ones . Owing to the noisy images and our findings we decided to add Label Smoothing and Cutout to training and it worked , we were able to achieve (lb 0.77 ) with the same ViT model . After that we tried with 448 Image size with the same strategy however it didn't work so well and hence was excluded from final ensemble . We trained another ViT model with different LR and increased Random Scale as augmentation and were able to scoer (lb 0.76)\n\nWe then started the search for diversifying our model zoo and moved on train SWIN and transformer-in-transformer with the same training Strategy , however Swin Beat tnt and scored (0.96 lb) as compared to tnt (1.07 lb)\n\n## RDkit Post-Processing\n\nEverybody was normalizing the predicitons and we were no different , but only after reading @nofreewill post we dug deeper and realized what we were missing.\n\n> def normalize_inchi(inchi):\n    try:\n        mol = Chem.MolFromInchi(inchi)\n    except:\n        pass\n    if mol is None:\n        return inchi,'invalid'\n    else:\n        try: \n            norm = Chem.MolToInchi(mol)\n            if norm == inchi:\n                return norm,'valid'\n            else:\n                return norm,'modified'\n        except: return inchi,'segmentation_failure'\n\nAfter adding this code to the normalization script we were able to get the status as well . We tested out the hypothesis that the places where INCHI were suggested invalid by rdkit were really wrong hence to replace them with valid predictions from different other model's prediction were a sure shot to improving score . We didn't stop here and dug even deeper to realize that the invalid INCHI's were coming from where the INCHI length was very long . Hence we fine-tuned our best tnt model on images having INCHi length greater than 150 and it gave great cv (0.99 previously 1.09) however due to lack of time we didn';t test lb and decided to add in post-processing to replace invalid INCHI's only\n\n## Models used in Post-Process to replace Invalid INCHI's\n\n* Tnt 336 Model \n* Vit 448 Model\n* Tnt 448 Model\n* prunedeffnetb3 + GRU Model\n* Effnetb4 + LSTM Model\n* Effnetb7 + GRU Model\n\n## Logits Ensemble\n\nI started coding Logit Ensemble on the last day and by the time it was ready only 9 hours were left in the competition , it gave a cv of 0.71 and hence we gave it our best to make it complete before 2 hours of deadline . Things fell into place and we were able to complete it in time.\n\nWe realized that in heng's code the final logit gave an output of <b>(bs,vocab_size) , so untill the max_length and the vocab size of models are same</b> we can at every time step ensemble the predictions , and since we didn't have a lot of time , we hard-coded it , something like this :\n\n> class EnsembleNet(nn.Module):\n    \n    def __init__(self, model_1, model_2, model_3):\n        super(EnsembleNet, self).__init__()\n        self.model_1 = model_1\n        self.model_2 = model_2\n        self.model_3 = model_3\n        \n        self.model_1.load_state_dict(torch.load(initial_checkpoint1)['state_dict'], strict=True)\n        self.model_2.load_state_dict(torch.load(initial_checkpoint2)['state_dict'], strict=True)\n        self.model_3.load_state_dict(torch.load(initial_checkpoint3)['state_dict'], strict=True)\n        \n        #self.model_1 = torch.jit.script(self.model_1)\n        #self.model_2 = torch.jit.script(self.model_2)\n        \n    def forward(self, image):\n        \n        image_dim = 768\n        text_dim = 768\n        decoder_dim = 768\n        num_layer = 4 #3\n        num_head = 8\n        ff_dim = 2048#1024\n\n        STOI = {\n            '<sos>': 190,\n            '<eos>': 191,\n            '<pad>': 192,\n            #'<mask>': 193,\n        }\n\n        image_size = 384\n        vocab_size = 193#194\n        max_length = 280 #300  # 275\n        \n        device = image.device\n        batch_size = len(image)\n        \n        image_embed_1 = self.model_1.cnn(image) #(bs,img_len,image_dim)\n        image_embed_1 = self.model_1.image_encode(image_embed_1).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        image_embed_2 = self.model_2.cnn(image) #(bs,img_len,image_dim)\n        image_embed_2 = self.model_2.image_encode(image_embed_2).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        image_embed_3 = self.model_3.cnn(image) #(bs,img_len,image_dim)\n        image_embed_3 = self.model_3.image_encode(image_embed_3).permute(1, 0, 2).contiguous() # (img_len,bs,image_dim)\n        \n        token = torch.full((batch_size, max_length), STOI['<pad>'], dtype=torch.long, device=device) # (batch_size,max_len) \n        text_pos_1 = self.model_1.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        text_pos_2 = self.model_2.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        text_pos_3 = self.model_3.text_pos.pos #(1,sequence_len,text_dim) torch.zeros(1, max_length, dim)\n        token[:, 0] = STOI['<sos>']\n        # -------------------------------------\n        eos = STOI['<eos>']\n        pad = STOI['<pad>']\n        # fast version\n        if 1:\n            # incremental_state = {}\n            incremental_state1 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            incremental_state2 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            incremental_state3 = torch.jit.annotate(\n                Dict[str, Dict[str, Optional[torch.Tensor]]],\n                torch.jit.annotate(Dict[str, Dict[str, Optional[torch.Tensor]]], {}),\n            )\n            for t in range(max_length - 1):\n                last_token_1 = token[:, t] # take the whole batch's t'th token\n                text_embed_1 = self.model_1.token_embed(last_token_1) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_1 = text_embed_1 + text_pos_1[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_1 = text_embed_1.reshape(1, batch_size, 768)\n                \n                last_token_2 = token[:, t] # take the whole batch's t'th token\n                text_embed_2 = self.model_2.token_embed(last_token_2) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_2 = text_embed_2 + text_pos_2[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_2 = text_embed_2.reshape(1, batch_size, 768)\n                \n                last_token_3 = token[:, t] # take the whole batch's t'th token\n                text_embed_3 = self.model_3.token_embed(last_token_3) #[bs,text_dim] Generate embedding for the t'th token\n                text_embed_3 = text_embed_3 + text_pos_3[:, t]  #[bs,text_dim] Combine with pos embed for t'th token\n                text_embed_3 = text_embed_3.reshape(1, batch_size, 1024)\n                \n                #text_embed ---> 1,bs,text_dim(768)\n                #image_embed ---> img_pos_embed,bs,image_im\n                x_1 = self.model_1.text_decode.forward_one(text_embed_1, image_embed_1, incremental_state1)\n                x_2 = self.model_2.text_decode.forward_one(text_embed_2, image_embed_2, incremental_state2)\n                x_3 = self.model_3.text_decode.forward_one(text_embed_3, image_embed_3, incremental_state3)\n                ## x -----> (1,bs,text_dim)\n                x_1 = x_1.reshape(batch_size, 768)\n                x_2 = x_2.reshape(batch_size, 768)\n                x_3 = x_3.reshape(batch_size, 1024)\n                ## x -----> (bs,decoder_dim)\n                l_1 = self.model_1.logit(x_1)\n                l_2 = self.model_2.logit(x_2)\n                l_3 = self.model_3.logit(x_3)\n                l = (l_1 + l_2 + l_3)/3\n                ## l -----> (bs,num_classes)\n                k = torch.argmax(l, -1)\n                token[:, t + 1] = k\n                if ((k == eos) | (k == pad)).all():\n                    break\n        \n        predict = token[:, 1:]\n        return predict\n\n## Final Hard-Voting\n\nOne Last Layer that we also added was hard-voting using all our submission (11 subs) . After doing RDkit Post Processing , we found that approx 11k invalid still remained and since we were not able to implement beam search for those invalid images we applied hard voting using all our model's submission ( @nischaydnk 's Idea)\n\n# Gratitude\n\nThis competition will remain closest to my heart , we have stayed up all night , its 6 AM in India right now , I will never forget how we passed the time from 1 to 4 while the ensemble inference was running . Congrats to @pheadrus @shivamcyborg for becoming a master and  @atsunorifujita @nischaydnk for their second gold . It was a collective team effort and I couldn;t thank you all enough for taking me in your team . \n\nIts late , I will clean up the code soon and post here  very soon . To be continued ....\nI hope you liked our solution . Thanks for reading",
    "1335032": "Congratulations on your gold medal @nischaydnk  @shivamcyborg 🎉🎉\n\nThe jump to the gold zone on the last day is amazing.😆",
    "1335034": "Congrats on you and your team's fist gold medal. Well done! :)\n@tanulsingh077 @atsunorifujita @shivamcyborg @pheadrus @nischaydnk",
    "1335037": "Congrats to competition master",
    "1335043": "Congrats for the gold :)",
    "1335064": "Congrats @tanulsingh077 . You did it bro, now competition master 🙌🙌🙌",
    "1335077": "Congrats @tanulsingh077 on the gold and becoming a competition master !! Well deserved!!",
    "1335093": "Congrats @tanulsingh077",
    "1335117": "Congrats @tanulsingh077",
    "1335226": "Congratulations @tanulsingh077 @nischaydnk @pheadrus @shivamcyborg @atsunorifujita for gold and becoming competition masters, an Amazing feat! Looking forward to learning from you guys 😄",
    "1335265": "Thank you @dehokanta , that was intense. 😄",
    "1335356": "Thanks a lot @heroseo , its time to move on to SETI I guess 😜",
    "1335358": "Thank you",
    "1335359": "Thanks a lot @Gao",
    "1335360": "Thank you @kamal , we finally got our gold",
    "1335361": "Thanks bro , it sure took long but finally it happened 😊",
    "1335362": "Thank you @atharva",
    "1335425": "Congratulation @tanulsingh077 for becoming competition Master and to the whole team for the outstanding results 🎉🎉. Interesting to notice how you improved the training of transformers models by reducing overfitting with cutout and label smoothing. We should have definitely tried that too. looking forward to seeing your code to learn more about your approach",
    "1335434": "Congratz to you and your team ! It was about time the \"hungry for gold\" team got its meal :)",
    "1335439": "Great work @tanulsingh077 and team! If I have to pick the thing that had the biggest impact it would be \"Label Smoothing and Cutout\" which got you from 1.27 to 0.77 LB. That's a HUGE boost. Would you agree? By the way what is this \"transformer to transformer\"? Are you referring to TNT?\n\nAgain, congrats on the standings in this competition, and for levelling up!",
    "1335481": "Congratz! Great job, you make it!",
    "1335483": "It always feels good when your idol appreciates you 😉 @theoviel , thanks a lot \nHungry for gold has a long way to go",
    "1335484": "Yes , that really surprised us even , but even more crucial was the logits ensemble which increased our best score without PP from 0.77 ----> 0.66 \n\nThe thing that helped us the most was constantly monitoring and analyzing the validation INCHI and trying to figure out what and where the models were doing wrong",
    "1335488": "Thanks a lot @jonathanbesomi , label smoothing was fujita's idea and we were also surprised of how good it worked",
    "1335511": "yes @alexandersoare by \"transformer to transformer\" we mean TNT and congratulations on your solo Silver medal 27th Place finish. 👍",
    "1335542": "Do you train your models in local machine or cloud?",
    "1335614": "We train models in cloud. Some of us have gpu locally setup in PC to run experiments.",
    "1335618": "Thanks @shivamcyborg !",
    "1335669": "Thanks @haqishen and many many congratulations to you and your team on gold 🥇 medal 5th place finish.",
    "1335672": "Thank you @dehokanta , last day idea worked pretty well here on time.",
    "1335681": "Thanks @heroseo",
    "1335696": "Hi may I know which cloud provider you guys used for your model training?",
    "1335730": "It was GCP and google colab pro from time to time.",
    "1335767": "Thanks a lot @haqishen",
    "1335883": "Tears, immense joy. Kaggle can bring out the best in us 😉\nCongrats on the last minute gold for you and team!",
    "1336114": "congrats for winning gold medal !!\nCould you please explain what you meant by transformer to transformer..? Is it transformer on encoder and decoder part..?\nif possible refer would be appreciated.",
    "1336484": "Congrats! Tears are a natural way of expressing the dedication and hardwork....More on your way",
    "1336881": "Thanks @rafiko1",
    "1336890": "Yes @prashantkramadhari transformer to transformer means transformer type encoder (like vision transformer , Swin or tnt ) and a transformer decoder (like vanilla transformer,BERT,etc)",
    "1336900": "Wow, Congrats Tanul!",
    "1337904": "Congratulations Mr_KnowNothing and team. Well done. I am happy that you got Gold. Great solution!",
    "1338100": "Thanks a lot @cdeotte , it feels great to be appreciated by you . I hope to continue this further",
    "1338102": "Congratulations!",
    "1338151": "Thank you @cdeotte :)",
    "1338570": "Congratulations, very well deserved!",
    "1534240": "Congratulations @tanulsingh077 @nischaydnk @pheadrus @shivamcyborg @atsunorifujita for your achievements."
  },
  "source": "meta"
}