{
  "id": 243787,
  "title": "4th Place Solution",
  "url": "/competitions/bms-molecular-translation/writeups/gpu-is-tired-4th-place-solution",
  "author_name": "",
  "post_date": "2021-06-05T03:35:30.267Z",
  "votes": 59,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Thanks to the kaggle and competition host for this competition. And also thanks to my teammate <a href=\"https://www.kaggle.com/haochensui\" target=\"_blank\">@haochensui</a>. We spent one month in this competition and in the first 2 weeks, we were doing experiments on 10% of the data. In the last two week, we rent several V100 and 3090 and switch to larger image size.</p>\n<p>Our solution is very simple and mainly composed of 4 models, plus some early models for post processing with rdkit. We use tokenizer from <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a>. Thanks for the amazing work.</p>\n<h4>Model Structures</h4>\n<p>CNN -&gt; Encoder -&gt; Decoder.<br>\nWe use resnest101 and efficientnetv2_m as our CNN backbones, and use 6 layers in transformer encoder and 9 layers in transformer decoder. We project embeddings from CNN into 512 channels to match the dimension of transformer.  In the encoder, we inject sinusoid positional embedding into each layer before key and query, pretty similar with DETR. In the decoder, we use relative positional embedding as in T5.</p>\n<h4>Model Training</h4>\n<p>Because of GPU limits, we did not train large scale models from scratch. We first train the model on 416x416 images and then increase the image resolutions and finetune the model. We reserve 2% data, ~48k for validation.</p>\n<p>We also do pseudo labeling here and find it quite useful. We used a 0.64CV/0.78LB submission(rdkit ensemble of efv2 on 416 and 640) to generate pseudo labels, and selected around 1.3M samples from it. We concatenate these samples with train samples. </p>\n<h6>Training order</h6>\n<p>efficientnetv2_m:  416(10 epochs) -&gt; 640(5 epochs)       -&gt;704(5 epochs + pseudo labels) </p>\n<p>​                                                                                               -&gt; 384x768 (5 epochs + pseudo labels)</p>\n<p>resnest101:            416(10 epochs)  -&gt; 640(5 epochs + pseudo labels)</p>\n<p>​                                                           -&gt;  384x768 (5 epochs + pseudo labels)  </p>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Size</th>\n<th>CV(no norm)</th>\n<th>LB(with norm)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2-m</td>\n<td>416</td>\n<td>0.99</td>\n<td>1.18</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>640</td>\n<td>0.74</td>\n<td>0.9</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>706</td>\n<td>0.65</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>384x768</td>\n<td>0.68</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>640</td>\n<td>0.71</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>384x768</td>\n<td>0.74</td>\n<td>0.67</td>\n</tr>\n</tbody>\n</table>\n<h4>Ensemble</h4>\n<p>We use 4 models with 0.66-0.67 LB to make a step-wise logit ensemble and use a beam search with beam size 3. Instead of choosing the prediction with maximum likelihood, we choose prediction with rdkit. We choose the first valid prediction from rdkit. In this way, we get around 0.02 improvement in CV. The ensemble of 4 models achieves 0.56 on LB. It should have  CV around 0.45. </p>\n<p>In addition to that, we merge our early submissions(scores &gt;0.65) with the 0.56 one according to rdkit and got 0.55.</p>\n<p>Finally, we select the 3500 invalid prediction in the submission and did a beam search with size 32. This reduces the LB score to 0.54.</p>",
  "messages": [
    {
      "id": "1335069",
      "postDate": "06/04/2021 02:12:19",
      "content": "<p>Thanks to the kaggle and competition host for this competition. And also thanks to my teammate <a href=\"https://www.kaggle.com/haochensui\" target=\"_blank\">@haochensui</a>. We spent one month in this competition and in the first 2 weeks, we were doing experiments on 10% of the data. In the last two week, we rent several V100 and 3090 and switch to larger image size.</p>\n<p>Our solution is very simple and mainly composed of 4 models, plus some early models for post processing with rdkit. We use tokenizer from <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a>. Thanks for the amazing work.</p>\n<h4>Model Structures</h4>\n<p>CNN -&gt; Encoder -&gt; Decoder.<br>\nWe use resnest101 and efficientnetv2_m as our CNN backbones, and use 6 layers in transformer encoder and 9 layers in transformer decoder. We project embeddings from CNN into 512 channels to match the dimension of transformer.  In the encoder, we inject sinusoid positional embedding into each layer before key and query, pretty similar with DETR. In the decoder, we use relative positional embedding as in T5.</p>\n<h4>Model Training</h4>\n<p>Because of GPU limits, we did not train large scale models from scratch. We first train the model on 416x416 images and then increase the image resolutions and finetune the model. We reserve 2% data, ~48k for validation.</p>\n<p>We also do pseudo labeling here and find it quite useful. We used a 0.64CV/0.78LB submission(rdkit ensemble of efv2 on 416 and 640) to generate pseudo labels, and selected around 1.3M samples from it. We concatenate these samples with train samples. </p>\n<h6>Training order</h6>\n<p>efficientnetv2_m:  416(10 epochs) -&gt; 640(5 epochs)       -&gt;704(5 epochs + pseudo labels) </p>\n<p>​                                                                                               -&gt; 384x768 (5 epochs + pseudo labels)</p>\n<p>resnest101:            416(10 epochs)  -&gt; 640(5 epochs + pseudo labels)</p>\n<p>​                                                           -&gt;  384x768 (5 epochs + pseudo labels)  </p>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Size</th>\n<th>CV(no norm)</th>\n<th>LB(with norm)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2-m</td>\n<td>416</td>\n<td>0.99</td>\n<td>1.18</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>640</td>\n<td>0.74</td>\n<td>0.9</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>706</td>\n<td>0.65</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>efficientnetv2-m</td>\n<td>384x768</td>\n<td>0.68</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>640</td>\n<td>0.71</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>384x768</td>\n<td>0.74</td>\n<td>0.67</td>\n</tr>\n</tbody>\n</table>\n<h4>Ensemble</h4>\n<p>We use 4 models with 0.66-0.67 LB to make a step-wise logit ensemble and use a beam search with beam size 3. Instead of choosing the prediction with maximum likelihood, we choose prediction with rdkit. We choose the first valid prediction from rdkit. In this way, we get around 0.02 improvement in CV. The ensemble of 4 models achieves 0.56 on LB. It should have  CV around 0.45. </p>\n<p>In addition to that, we merge our early submissions(scores &gt;0.65) with the 0.56 one according to rdkit and got 0.55.</p>\n<p>Finally, we select the 3500 invalid prediction in the submission and did a beam search with size 32. This reduces the LB score to 0.54.</p>",
      "rawMarkdown": "Thanks to the kaggle and competition host for this competition. And also thanks to my teammate @haochensui. We spent one month in this competition and in the first 2 weeks, we were doing experiments on 10% of the data. In the last two week, we rent several V100 and 3090 and switch to larger image size.\n\nOur solution is very simple and mainly composed of 4 models, plus some early models for post processing with rdkit. We use tokenizer from https://www.kaggle.com/yasufuminakama/inchi-preprocess-2. Thanks for the amazing work.\n\n#### Model Structures \n\n\nCNN -> Encoder -> Decoder.\nWe use resnest101 and efficientnetv2_m as our CNN backbones, and use 6 layers in transformer encoder and 9 layers in transformer decoder. We project embeddings from CNN into 512 channels to match the dimension of transformer.  In the encoder, we inject sinusoid positional embedding into each layer before key and query, pretty similar with DETR. In the decoder, we use relative positional embedding as in T5.\n\n#### Model Training\n\nBecause of GPU limits, we did not train large scale models from scratch. We first train the model on 416x416 images and then increase the image resolutions and finetune the model. We reserve 2% data, ~48k for validation.\n\nWe also do pseudo labeling here and find it quite useful. We used a 0.64CV/0.78LB submission(rdkit ensemble of efv2 on 416 and 640) to generate pseudo labels, and selected around 1.3M samples from it. We concatenate these samples with train samples. \n\n###### Training order\n\nefficientnetv2_m:  416(10 epochs) -> 640(5 epochs)       ->704(5 epochs + pseudo labels) \n\n​                                                                                               -> 384x768 (5 epochs + pseudo labels)\n\nresnest101:            416(10 epochs)  -> 640(5 epochs + pseudo labels)\n\n​                                                           ->  384x768 (5 epochs + pseudo labels)  \n\n#### Results\n\n| Model            | Size    | CV(no norm) | LB(with norm) |\n| ---------------- | ------- | ----------- | ------------- |\n| efficientnetv2-m | 416     | 0.99        | 1.18          |\n| efficientnetv2-m | 640     | 0.74        | 0.9           |\n| efficientnetv2-m | 706     | 0.65        | 0.67          |\n| efficientnetv2-m | 384x768 | 0.68        | 0.67          |\n| resnest101       | 640     | 0.71        | 0.66          |\n| resnest101       | 384x768 | 0.74        | 0.67          |\n\n#### Ensemble\n\nWe use 4 models with 0.66-0.67 LB to make a step-wise logit ensemble and use a beam search with beam size 3. Instead of choosing the prediction with maximum likelihood, we choose prediction with rdkit. We choose the first valid prediction from rdkit. In this way, we get around 0.02 improvement in CV. The ensemble of 4 models achieves 0.56 on LB. It should have  CV around 0.45. \n\nIn addition to that, we merge our early submissions(scores >0.65) with the 0.56 one according to rdkit and got 0.55.\n\nFinally, we select the 3500 invalid prediction in the submission and did a beam search with size 32. This reduces the LB score to 0.54.",
      "votes": null
    },
    {
      "id": "1335070",
      "postDate": "06/04/2021 02:13:27",
      "content": "<p>Congrats GM!</p>",
      "rawMarkdown": "Congrats GM!",
      "votes": null
    },
    {
      "id": "1335071",
      "postDate": "06/04/2021 02:16:02",
      "content": "<p>Thanks, your tokenizer saved me a lot of work</p>",
      "rawMarkdown": "Thanks, your tokenizer saved me a lot of work",
      "votes": null
    },
    {
      "id": "1335076",
      "postDate": "06/04/2021 02:28:33",
      "content": "<p>I am glad you finally become GM and nice write up =) </p>",
      "rawMarkdown": "I am glad you finally become GM and nice write up =)",
      "votes": null
    },
    {
      "id": "1335141",
      "postDate": "06/04/2021 03:17:28",
      "content": "<p>thanks for sharing.  I noticed  normalize_inchi() and rdkit in discussion. Due to failing to install rdkit on colab, I gave up.<br>\nprobably It could improve my LB a little bit if I used it?</p>",
      "rawMarkdown": "thanks for sharing.  I noticed  normalize_inchi() and rdkit in discussion. Due to failing to install rdkit on colab, I gave up.\nprobably It could improve my LB a little bit if I used it?",
      "votes": null
    },
    {
      "id": "1335154",
      "postDate": "06/04/2021 03:27:50",
      "content": "<p>pip install rdkit-pypi<br>\nI use this one to install </p>",
      "rawMarkdown": "pip install rdkit-pypi\nI use this one to install",
      "votes": null
    },
    {
      "id": "1335340",
      "postDate": "06/04/2021 06:51:07",
      "content": "<p>Congrats to you and your team! New GM!</p>",
      "rawMarkdown": "Congrats to you and your team! New GM!",
      "votes": null
    },
    {
      "id": "1335349",
      "postDate": "06/04/2021 06:56:53",
      "content": "<p>Thanks for the write-up , congratulations on becoming a GM , can you please share the code with us especially for beam search , I was not able to make it work with transformers , there is still a lot to learn</p>",
      "rawMarkdown": "Thanks for the write-up , congratulations on becoming a GM , can you please share the code with us especially for beam search , I was not able to make it work with transformers , there is still a lot to learn",
      "votes": null
    },
    {
      "id": "1335831",
      "postDate": "06/04/2021 13:12:03",
      "content": "<p>Congratulations ! thanks for sharing your approach </p>",
      "rawMarkdown": "Congratulations ! thanks for sharing your approach",
      "votes": null
    },
    {
      "id": "1335842",
      "postDate": "06/04/2021 13:22:08",
      "content": "<p>Nice progress in the last week, congrats!</p>",
      "rawMarkdown": "Nice progress in the last week, congrats!",
      "votes": null
    },
    {
      "id": "1336012",
      "postDate": "06/04/2021 15:28:58",
      "content": "<p>Congratz to our new GMs!</p>",
      "rawMarkdown": "Congratz to our new GMs!",
      "votes": null
    },
    {
      "id": "1336433",
      "postDate": "06/04/2021 22:38:37",
      "content": "<p>I use this class to track the scores, sequences and caches of the transformer.</p>\n<pre><code>class BeamSearcher_Transformer_Ensemble():\n    def __init__(self,batch_size,beam_size,sos_idx,eos_idx,device):\n        self.beam_size=beam_size\n        self.batch_size=batch_size\n        self.eos_idx=eos_idx\n        self.sos_idx=sos_idx\n        self.device=device\n        self.current_sequence,self.current_cache=self.init_input(batch_size)\n        self.current_scores= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n        self.finish_mask= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device).bool()\n        self.sentence_length= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n\n    def init_input(self,batch_size):\n        init_decoded_tokens = torch.LongTensor([self.sos_idx]).to(self.device).reshape(1,1,1).repeat(1,batch_size,self.beam_size) # [1, bs,beam_size]\n        init_cache = None\n        return init_decoded_tokens,init_cache\n\n    def get_input(self):\n        #current_sequence is a [T,bs,beam_size] LongTensor\n        return self.current_sequence,self.current_cache\n\n    def rearrange_cache(self,beam_group,cache):\n        '''\n        Args:\n            beam_group: [batch_size,beam_size]\n            cache:      list[batch_size*beam_size,T,hdim]\n\n        Returns:\n            list[batch_size*beam_size,T,hdim]\n        '''\n        new_cache=[]\n        batch_idx = torch.arange(self.batch_size, device=cache[0].device)\n        for c in cache:\n            c=c.unflatten(0,[self.batch_size,self.beam_size])\n            new_c=torch.stack([c[batch_idx, beam_group[:, x]] for x in range(beam_group.shape[1])], dim=1)\n            new_cache.append(new_c.flatten(0,1))\n        return new_cache\n\n    def prepare_next_input(self,predictions,new_cache,first_step):\n        '''\n        Args:\n            predictions: [bs*beam_size,V] logits\n            new_cache: list[list[bs*beam_size,T,Hidden]]\n            first_step:  if True, this is the first step, a sentence is copied for n times, we need to mask out copies\n        Returns:\n            finished: if True, all sequences is finished\n        '''\n        V=predictions.shape[-1]\n        predictions=predictions.view(self.batch_size,self.beam_size,V)\n        loglikelihood=torch.log_softmax(predictions,-1) #[bs,beam_size,V]\n\n        if first_step:\n            loglikelihood[:,1:]=-float('inf')\n\n        loglikelihood[self.finish_mask]=-float('inf')\n        finished=torch.nonzero(self.finish_mask,as_tuple=True)\n        loglikelihood[(finished[0],finished[1],torch.ones_like(finished[0])*self.eos_idx)]=0.\n\n        new_scores= self.current_scores.unsqueeze(-1)+loglikelihood #[bs,beam_size,V]\n        new_length= self.sentence_length + 1-self.finish_mask.float() #[bs,beam_size] descible length of sentence\n\n        normalized_scores= (new_scores/new_length.unsqueeze(-1)).view(self.batch_size,self.beam_size*V) #[bs,beam_size*V]\n        #normalized_scores = (new_scores).view(self.batch_size,self.beam_size * V)  # [bs,beam_size*V]\n        _,top_idx=torch.topk(normalized_scores,dim=-1,largest=True,k=self.beam_size)\n        beam_group=top_idx//V #[bs,beam_size]\n        token_idx=top_idx%V  #[bs,beam_size] #index of new tokens\n\n        #update scores, mask, length\n        self.sentence_length=beam_index_select(new_length,beam_group)\n        self.current_scores=torch.gather(new_scores.view(self.batch_size,self.beam_size*V),dim=1,index=top_idx)\n        self.finish_mask=torch.logical_or(beam_index_select(self.finish_mask,beam_group),token_idx==self.eos_idx)\n        self.current_sequence= beam_index_select(self.current_sequence.permute(1,2,0),beam_group).permute(2,0,1)\n\n        self.current_sequence=torch.cat([self.current_sequence,token_idx.unsqueeze(0)],dim=0)\n        self.current_cache=[self.rearrange_cache(beam_group,x) for x in new_cache]\n        return torch.all(self.finish_mask)\n</code></pre>",
      "rawMarkdown": "I use this class to track the scores, sequences and caches of the transformer.\n```\nclass BeamSearcher_Transformer_Ensemble():\n    def __init__(self,batch_size,beam_size,sos_idx,eos_idx,device):\n        self.beam_size=beam_size\n        self.batch_size=batch_size\n        self.eos_idx=eos_idx\n        self.sos_idx=sos_idx\n        self.device=device\n        self.current_sequence,self.current_cache=self.init_input(batch_size)\n        self.current_scores= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n        self.finish_mask= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device).bool()\n        self.sentence_length= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n\n    def init_input(self,batch_size):\n        init_decoded_tokens = torch.LongTensor([self.sos_idx]).to(self.device).reshape(1,1,1).repeat(1,batch_size,self.beam_size) # [1, bs,beam_size]\n        init_cache = None\n        return init_decoded_tokens,init_cache\n\n    def get_input(self):\n        #current_sequence is a [T,bs,beam_size] LongTensor\n        return self.current_sequence,self.current_cache\n\n    def rearrange_cache(self,beam_group,cache):\n        '''\n        Args:\n            beam_group: [batch_size,beam_size]\n            cache:      list[batch_size*beam_size,T,hdim]\n\n        Returns:\n            list[batch_size*beam_size,T,hdim]\n        '''\n        new_cache=[]\n        batch_idx = torch.arange(self.batch_size, device=cache[0].device)\n        for c in cache:\n            c=c.unflatten(0,[self.batch_size,self.beam_size])\n            new_c=torch.stack([c[batch_idx, beam_group[:, x]] for x in range(beam_group.shape[1])], dim=1)\n            new_cache.append(new_c.flatten(0,1))\n        return new_cache\n\n    def prepare_next_input(self,predictions,new_cache,first_step):\n        '''\n        Args:\n            predictions: [bs*beam_size,V] logits\n            new_cache: list[list[bs*beam_size,T,Hidden]]\n            first_step:  if True, this is the first step, a sentence is copied for n times, we need to mask out copies\n        Returns:\n            finished: if True, all sequences is finished\n        '''\n        V=predictions.shape[-1]\n        predictions=predictions.view(self.batch_size,self.beam_size,V)\n        loglikelihood=torch.log_softmax(predictions,-1) #[bs,beam_size,V]\n\n        if first_step:\n            loglikelihood[:,1:]=-float('inf')\n\n        loglikelihood[self.finish_mask]=-float('inf')\n        finished=torch.nonzero(self.finish_mask,as_tuple=True)\n        loglikelihood[(finished[0],finished[1],torch.ones_like(finished[0])*self.eos_idx)]=0.\n\n        new_scores= self.current_scores.unsqueeze(-1)+loglikelihood #[bs,beam_size,V]\n        new_length= self.sentence_length + 1-self.finish_mask.float() #[bs,beam_size] descible length of sentence\n\n        normalized_scores= (new_scores/new_length.unsqueeze(-1)).view(self.batch_size,self.beam_size*V) #[bs,beam_size*V]\n        #normalized_scores = (new_scores).view(self.batch_size,self.beam_size * V)  # [bs,beam_size*V]\n        _,top_idx=torch.topk(normalized_scores,dim=-1,largest=True,k=self.beam_size)\n        beam_group=top_idx//V #[bs,beam_size]\n        token_idx=top_idx%V  #[bs,beam_size] #index of new tokens\n\n        #update scores, mask, length\n        self.sentence_length=beam_index_select(new_length,beam_group)\n        self.current_scores=torch.gather(new_scores.view(self.batch_size,self.beam_size*V),dim=1,index=top_idx)\n        self.finish_mask=torch.logical_or(beam_index_select(self.finish_mask,beam_group),token_idx==self.eos_idx)\n        self.current_sequence= beam_index_select(self.current_sequence.permute(1,2,0),beam_group).permute(2,0,1)\n\n        self.current_sequence=torch.cat([self.current_sequence,token_idx.unsqueeze(0)],dim=0)\n        self.current_cache=[self.rearrange_cache(beam_group,x) for x in new_cache]\n        return torch.all(self.finish_mask)\n```",
      "votes": null
    },
    {
      "id": "1336563",
      "postDate": "06/05/2021 03:32:43",
      "content": "<p>Cangrats to be a GM!<br>\nYour single models are so strong themselves! Can you give me a hint about progressive training in image size?<br>\nI've tried same but didn't work due to positional encoding. How did you put multiple size pos enc to encoder output? (I also add pos enc to only k, v before attention)</p>",
      "rawMarkdown": "Cangrats to be a GM!\nYour single models are so strong themselves! Can you give me a hint about progressive training in image size?\nI've tried same but didn't work due to positional encoding. How did you put multiple size pos enc to encoder output? (I also add pos enc to only k, v before attention)",
      "votes": null
    },
    {
      "id": "1336566",
      "postDate": "06/05/2021 03:40:25",
      "content": "<p>Oh, the white space is not displayed well here. <br>\nFor efficientnetv2, I train from 416 to 640 and then I finetune it on two sizes,  704 and 384x768.<br>\nFor resnest101, I trained 416 and then finetune on two sizes, 640 and 384x768</p>",
      "rawMarkdown": "Oh, the white space is not displayed well here. \nFor efficientnetv2, I train from 416 to 640 and then I finetune it on two sizes,  704 and 384x768.\nFor resnest101, I trained 416 and then finetune on two sizes, 640 and 384x768",
      "votes": null
    },
    {
      "id": "1336638",
      "postDate": "06/05/2021 05:28:29",
      "content": "<p>But I just change to a new size and tune the model for 5 more epochs. The loss went down very quickly despite the encoding changed. <br>\nPositional encodings on CNN output has small effects on the results, since CNN can learn positional information from zero padding. </p>",
      "rawMarkdown": "But I just change to a new size and tune the model for 5 more epochs. The loss went down very quickly despite the encoding changed. \nPositional encodings on CNN output has small effects on the results, since CNN can learn positional information from zero padding.",
      "votes": null
    },
    {
      "id": "1336660",
      "postDate": "06/05/2021 05:51:36",
      "content": "<p>thanks for sharing.  I used conda installation searched from online, some problem, plus other issues, I gave up.</p>\n<p>my CV around 1.1 but LB 2.35.   Training without noise or augmentation?    leaders CV and LB are very close.</p>",
      "rawMarkdown": "thanks for sharing.  I used conda installation searched from online, some problem, plus other issues, I gave up.\n\nmy CV around 1.1 but LB 2.35.   Training without noise or augmentation?    leaders CV and LB are very close.",
      "votes": null
    },
    {
      "id": "1336667",
      "postDate": "06/05/2021 06:03:31",
      "content": "<p>spent money on renting hardware resources. haha.</p>",
      "rawMarkdown": "spent money on renting hardware resources. haha.",
      "votes": null
    },
    {
      "id": "1336880",
      "postDate": "06/05/2021 09:05:37",
      "content": "<p>Thanks for the code </p>",
      "rawMarkdown": "Thanks for the code",
      "votes": null
    },
    {
      "id": "1337191",
      "postDate": "06/05/2021 13:20:31",
      "content": "<p>Congratulations on making it to GM! Can I ask - how did you get around the rdkit seg fault issue? I thought about using this in conjunction with beam search but I was worried that the code would just keep falling over.</p>",
      "rawMarkdown": "Congratulations on making it to GM! Can I ask - how did you get around the rdkit seg fault issue? I thought about using this in conjunction with beam search but I was worried that the code would just keep falling over.",
      "votes": null
    },
    {
      "id": "1337243",
      "postDate": "06/05/2021 13:51:49",
      "content": "<p>I'm not the author, but I used multiprocessing Pool, that automatically restarts failed child processes:</p>\n<pre><code>from multiprocessing import Pool\nfrom multiprocessing.context import TimeoutError\n\ndef normalize_inchi(inchi):\n    try:\n        return Chem.MolToInchi(Chem.MolFromInchi(inchi))\n    except:\n        return None\n\npool = Pool(1)\nr = pool.apply_async(normalize_inchi, (inchi,))\ntry:\n    normalized_inchi = r.get(timeout=1)\nexcept TimeoutError:\n    normalized_inchi = None\n</code></pre>",
      "rawMarkdown": "I'm not the author, but I used multiprocessing Pool, that automatically restarts failed child processes:\n```\nfrom multiprocessing import Pool\nfrom multiprocessing.context import TimeoutError\n\ndef normalize_inchi(inchi):\n    try:\n        return Chem.MolToInchi(Chem.MolFromInchi(inchi))\n    except:\n        return None\n\npool = Pool(1)\nr = pool.apply_async(normalize_inchi, (inchi,))\ntry:\n    normalized_inchi = r.get(timeout=1)\nexcept TimeoutError:\n    normalized_inchi = None\n```",
      "votes": null
    },
    {
      "id": "1337404",
      "postDate": "06/05/2021 15:37:41",
      "content": "<p>Hey, are you going to share the notebook for this cause I wanna check out the code.</p>",
      "rawMarkdown": "Hey, are you going to share the notebook for this cause I wanna check out the code.",
      "votes": null
    },
    {
      "id": "1337490",
      "postDate": "06/05/2021 16:29:06",
      "content": "<p><a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> Ah yes, of course, clever! That's the way to do it. I had been doing the full test set using a bash script - checking the exit code and continuing to loop if it wasn't valid. </p>",
      "rawMarkdown": "stassl Ah yes, of course, clever! That's the way to do it. I had been doing the full test set using a bash script - checking the exit code and continuing to loop if it wasn't valid.",
      "votes": null
    },
    {
      "id": "1337689",
      "postDate": "06/05/2021 19:31:12",
      "content": "<p>That's a nice solution, but in my case there is only 0-2 segmentation faults. So I just mannually rerun the code when it fails.</p>",
      "rawMarkdown": "That's a nice solution, but in my case there is only 0-2 segmentation faults. So I just mannually rerun the code when it fails.",
      "votes": null
    },
    {
      "id": "1340781",
      "postDate": "06/08/2021 08:58:15",
      "content": "<p>Congratulations.</p>",
      "rawMarkdown": "Congratulations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335070,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "06/04/2021 02:13:27",
      "content": "<p>Congrats GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335071,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/04/2021 02:16:02",
          "content": "<p>Thanks, your tokenizer saved me a lot of work</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335076,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "06/04/2021 02:28:33",
      "content": "<p>I am glad you finally become GM and nice write up =) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335141,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/04/2021 03:17:28",
      "content": "<p>thanks for sharing.  I noticed  normalize_inchi() and rdkit in discussion. Due to failing to install rdkit on colab, I gave up.<br>\nprobably It could improve my LB a little bit if I used it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335154,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/04/2021 03:27:50",
          "content": "<p>pip install rdkit-pypi<br>\nI use this one to install </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336660,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "06/05/2021 05:51:36",
          "content": "<p>thanks for sharing.  I used conda installation searched from online, some problem, plus other issues, I gave up.</p>\n<p>my CV around 1.1 but LB 2.35.   Training without noise or augmentation?    leaders CV and LB are very close.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335340,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "06/04/2021 06:51:07",
      "content": "<p>Congrats to you and your team! New GM!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335349,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "06/04/2021 06:56:53",
      "content": "<p>Thanks for the write-up , congratulations on becoming a GM , can you please share the code with us especially for beam search , I was not able to make it work with transformers , there is still a lot to learn</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336433,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/04/2021 22:38:37",
          "content": "<p>I use this class to track the scores, sequences and caches of the transformer.</p>\n<pre><code>class BeamSearcher_Transformer_Ensemble():\n    def __init__(self,batch_size,beam_size,sos_idx,eos_idx,device):\n        self.beam_size=beam_size\n        self.batch_size=batch_size\n        self.eos_idx=eos_idx\n        self.sos_idx=sos_idx\n        self.device=device\n        self.current_sequence,self.current_cache=self.init_input(batch_size)\n        self.current_scores= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n        self.finish_mask= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device).bool()\n        self.sentence_length= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n\n    def init_input(self,batch_size):\n        init_decoded_tokens = torch.LongTensor([self.sos_idx]).to(self.device).reshape(1,1,1).repeat(1,batch_size,self.beam_size) # [1, bs,beam_size]\n        init_cache = None\n        return init_decoded_tokens,init_cache\n\n    def get_input(self):\n        #current_sequence is a [T,bs,beam_size] LongTensor\n        return self.current_sequence,self.current_cache\n\n    def rearrange_cache(self,beam_group,cache):\n        '''\n        Args:\n            beam_group: [batch_size,beam_size]\n            cache:      list[batch_size*beam_size,T,hdim]\n\n        Returns:\n            list[batch_size*beam_size,T,hdim]\n        '''\n        new_cache=[]\n        batch_idx = torch.arange(self.batch_size, device=cache[0].device)\n        for c in cache:\n            c=c.unflatten(0,[self.batch_size,self.beam_size])\n            new_c=torch.stack([c[batch_idx, beam_group[:, x]] for x in range(beam_group.shape[1])], dim=1)\n            new_cache.append(new_c.flatten(0,1))\n        return new_cache\n\n    def prepare_next_input(self,predictions,new_cache,first_step):\n        '''\n        Args:\n            predictions: [bs*beam_size,V] logits\n            new_cache: list[list[bs*beam_size,T,Hidden]]\n            first_step:  if True, this is the first step, a sentence is copied for n times, we need to mask out copies\n        Returns:\n            finished: if True, all sequences is finished\n        '''\n        V=predictions.shape[-1]\n        predictions=predictions.view(self.batch_size,self.beam_size,V)\n        loglikelihood=torch.log_softmax(predictions,-1) #[bs,beam_size,V]\n\n        if first_step:\n            loglikelihood[:,1:]=-float('inf')\n\n        loglikelihood[self.finish_mask]=-float('inf')\n        finished=torch.nonzero(self.finish_mask,as_tuple=True)\n        loglikelihood[(finished[0],finished[1],torch.ones_like(finished[0])*self.eos_idx)]=0.\n\n        new_scores= self.current_scores.unsqueeze(-1)+loglikelihood #[bs,beam_size,V]\n        new_length= self.sentence_length + 1-self.finish_mask.float() #[bs,beam_size] descible length of sentence\n\n        normalized_scores= (new_scores/new_length.unsqueeze(-1)).view(self.batch_size,self.beam_size*V) #[bs,beam_size*V]\n        #normalized_scores = (new_scores).view(self.batch_size,self.beam_size * V)  # [bs,beam_size*V]\n        _,top_idx=torch.topk(normalized_scores,dim=-1,largest=True,k=self.beam_size)\n        beam_group=top_idx//V #[bs,beam_size]\n        token_idx=top_idx%V  #[bs,beam_size] #index of new tokens\n\n        #update scores, mask, length\n        self.sentence_length=beam_index_select(new_length,beam_group)\n        self.current_scores=torch.gather(new_scores.view(self.batch_size,self.beam_size*V),dim=1,index=top_idx)\n        self.finish_mask=torch.logical_or(beam_index_select(self.finish_mask,beam_group),token_idx==self.eos_idx)\n        self.current_sequence= beam_index_select(self.current_sequence.permute(1,2,0),beam_group).permute(2,0,1)\n\n        self.current_sequence=torch.cat([self.current_sequence,token_idx.unsqueeze(0)],dim=0)\n        self.current_cache=[self.rearrange_cache(beam_group,x) for x in new_cache]\n        return torch.all(self.finish_mask)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336880,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/05/2021 09:05:37",
          "content": "<p>Thanks for the code </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335831,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "06/04/2021 13:12:03",
      "content": "<p>Congratulations ! thanks for sharing your approach </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335842,
      "author_name": "stassl",
      "author_url": "",
      "post_date": "06/04/2021 13:22:08",
      "content": "<p>Nice progress in the last week, congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336667,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "06/05/2021 06:03:31",
          "content": "<p>spent money on renting hardware resources. haha.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336012,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "06/04/2021 15:28:58",
      "content": "<p>Congratz to our new GMs!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336563,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "06/05/2021 03:32:43",
      "content": "<p>Cangrats to be a GM!<br>\nYour single models are so strong themselves! Can you give me a hint about progressive training in image size?<br>\nI've tried same but didn't work due to positional encoding. How did you put multiple size pos enc to encoder output? (I also add pos enc to only k, v before attention)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336566,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/05/2021 03:40:25",
          "content": "<p>Oh, the white space is not displayed well here. <br>\nFor efficientnetv2, I train from 416 to 640 and then I finetune it on two sizes,  704 and 384x768.<br>\nFor resnest101, I trained 416 and then finetune on two sizes, 640 and 384x768</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336638,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/05/2021 05:28:29",
          "content": "<p>But I just change to a new size and tune the model for 5 more epochs. The loss went down very quickly despite the encoding changed. <br>\nPositional encodings on CNN output has small effects on the results, since CNN can learn positional information from zero padding. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337191,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/05/2021 13:20:31",
      "content": "<p>Congratulations on making it to GM! Can I ask - how did you get around the rdkit seg fault issue? I thought about using this in conjunction with beam search but I was worried that the code would just keep falling over.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337243,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "06/05/2021 13:51:49",
          "content": "<p>I'm not the author, but I used multiprocessing Pool, that automatically restarts failed child processes:</p>\n<pre><code>from multiprocessing import Pool\nfrom multiprocessing.context import TimeoutError\n\ndef normalize_inchi(inchi):\n    try:\n        return Chem.MolToInchi(Chem.MolFromInchi(inchi))\n    except:\n        return None\n\npool = Pool(1)\nr = pool.apply_async(normalize_inchi, (inchi,))\ntry:\n    normalized_inchi = r.get(timeout=1)\nexcept TimeoutError:\n    normalized_inchi = None\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337490,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "06/05/2021 16:29:06",
          "content": "<p><a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> Ah yes, of course, clever! That's the way to do it. I had been doing the full test set using a bash script - checking the exit code and continuing to loop if it wasn't valid. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337689,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "06/05/2021 19:31:12",
          "content": "<p>That's a nice solution, but in my case there is only 0-2 segmentation faults. So I just mannually rerun the code when it fails.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337404,
      "author_name": "bamblebam",
      "author_url": "",
      "post_date": "06/05/2021 15:37:41",
      "content": "<p>Hey, are you going to share the notebook for this cause I wanna check out the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340781,
      "author_name": "sohailds",
      "author_url": "",
      "post_date": "06/08/2021 08:58:15",
      "content": "<p>Congratulations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335069": "Thanks to the kaggle and competition host for this competition. And also thanks to my teammate @haochensui. We spent one month in this competition and in the first 2 weeks, we were doing experiments on 10% of the data. In the last two week, we rent several V100 and 3090 and switch to larger image size.\n\nOur solution is very simple and mainly composed of 4 models, plus some early models for post processing with rdkit. We use tokenizer from https://www.kaggle.com/yasufuminakama/inchi-preprocess-2. Thanks for the amazing work.\n\n#### Model Structures \n\n\nCNN -> Encoder -> Decoder.\nWe use resnest101 and efficientnetv2_m as our CNN backbones, and use 6 layers in transformer encoder and 9 layers in transformer decoder. We project embeddings from CNN into 512 channels to match the dimension of transformer.  In the encoder, we inject sinusoid positional embedding into each layer before key and query, pretty similar with DETR. In the decoder, we use relative positional embedding as in T5.\n\n#### Model Training\n\nBecause of GPU limits, we did not train large scale models from scratch. We first train the model on 416x416 images and then increase the image resolutions and finetune the model. We reserve 2% data, ~48k for validation.\n\nWe also do pseudo labeling here and find it quite useful. We used a 0.64CV/0.78LB submission(rdkit ensemble of efv2 on 416 and 640) to generate pseudo labels, and selected around 1.3M samples from it. We concatenate these samples with train samples. \n\n###### Training order\n\nefficientnetv2_m:  416(10 epochs) -> 640(5 epochs)       ->704(5 epochs + pseudo labels) \n\n​                                                                                               -> 384x768 (5 epochs + pseudo labels)\n\nresnest101:            416(10 epochs)  -> 640(5 epochs + pseudo labels)\n\n​                                                           ->  384x768 (5 epochs + pseudo labels)  \n\n#### Results\n\n| Model            | Size    | CV(no norm) | LB(with norm) |\n| ---------------- | ------- | ----------- | ------------- |\n| efficientnetv2-m | 416     | 0.99        | 1.18          |\n| efficientnetv2-m | 640     | 0.74        | 0.9           |\n| efficientnetv2-m | 706     | 0.65        | 0.67          |\n| efficientnetv2-m | 384x768 | 0.68        | 0.67          |\n| resnest101       | 640     | 0.71        | 0.66          |\n| resnest101       | 384x768 | 0.74        | 0.67          |\n\n#### Ensemble\n\nWe use 4 models with 0.66-0.67 LB to make a step-wise logit ensemble and use a beam search with beam size 3. Instead of choosing the prediction with maximum likelihood, we choose prediction with rdkit. We choose the first valid prediction from rdkit. In this way, we get around 0.02 improvement in CV. The ensemble of 4 models achieves 0.56 on LB. It should have  CV around 0.45. \n\nIn addition to that, we merge our early submissions(scores >0.65) with the 0.56 one according to rdkit and got 0.55.\n\nFinally, we select the 3500 invalid prediction in the submission and did a beam search with size 32. This reduces the LB score to 0.54.",
    "1335070": "Congrats GM!",
    "1335071": "Thanks, your tokenizer saved me a lot of work",
    "1335076": "I am glad you finally become GM and nice write up =)",
    "1335141": "thanks for sharing.  I noticed  normalize_inchi() and rdkit in discussion. Due to failing to install rdkit on colab, I gave up.\nprobably It could improve my LB a little bit if I used it?",
    "1335154": "pip install rdkit-pypi\nI use this one to install",
    "1335340": "Congrats to you and your team! New GM!",
    "1335349": "Thanks for the write-up , congratulations on becoming a GM , can you please share the code with us especially for beam search , I was not able to make it work with transformers , there is still a lot to learn",
    "1335831": "Congratulations ! thanks for sharing your approach",
    "1335842": "Nice progress in the last week, congrats!",
    "1336012": "Congratz to our new GMs!",
    "1336433": "I use this class to track the scores, sequences and caches of the transformer.\n```\nclass BeamSearcher_Transformer_Ensemble():\n    def __init__(self,batch_size,beam_size,sos_idx,eos_idx,device):\n        self.beam_size=beam_size\n        self.batch_size=batch_size\n        self.eos_idx=eos_idx\n        self.sos_idx=sos_idx\n        self.device=device\n        self.current_sequence,self.current_cache=self.init_input(batch_size)\n        self.current_scores= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n        self.finish_mask= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device).bool()\n        self.sentence_length= torch.zeros([batch_size,beam_size],dtype=torch.float32,device=self.device)\n\n    def init_input(self,batch_size):\n        init_decoded_tokens = torch.LongTensor([self.sos_idx]).to(self.device).reshape(1,1,1).repeat(1,batch_size,self.beam_size) # [1, bs,beam_size]\n        init_cache = None\n        return init_decoded_tokens,init_cache\n\n    def get_input(self):\n        #current_sequence is a [T,bs,beam_size] LongTensor\n        return self.current_sequence,self.current_cache\n\n    def rearrange_cache(self,beam_group,cache):\n        '''\n        Args:\n            beam_group: [batch_size,beam_size]\n            cache:      list[batch_size*beam_size,T,hdim]\n\n        Returns:\n            list[batch_size*beam_size,T,hdim]\n        '''\n        new_cache=[]\n        batch_idx = torch.arange(self.batch_size, device=cache[0].device)\n        for c in cache:\n            c=c.unflatten(0,[self.batch_size,self.beam_size])\n            new_c=torch.stack([c[batch_idx, beam_group[:, x]] for x in range(beam_group.shape[1])], dim=1)\n            new_cache.append(new_c.flatten(0,1))\n        return new_cache\n\n    def prepare_next_input(self,predictions,new_cache,first_step):\n        '''\n        Args:\n            predictions: [bs*beam_size,V] logits\n            new_cache: list[list[bs*beam_size,T,Hidden]]\n            first_step:  if True, this is the first step, a sentence is copied for n times, we need to mask out copies\n        Returns:\n            finished: if True, all sequences is finished\n        '''\n        V=predictions.shape[-1]\n        predictions=predictions.view(self.batch_size,self.beam_size,V)\n        loglikelihood=torch.log_softmax(predictions,-1) #[bs,beam_size,V]\n\n        if first_step:\n            loglikelihood[:,1:]=-float('inf')\n\n        loglikelihood[self.finish_mask]=-float('inf')\n        finished=torch.nonzero(self.finish_mask,as_tuple=True)\n        loglikelihood[(finished[0],finished[1],torch.ones_like(finished[0])*self.eos_idx)]=0.\n\n        new_scores= self.current_scores.unsqueeze(-1)+loglikelihood #[bs,beam_size,V]\n        new_length= self.sentence_length + 1-self.finish_mask.float() #[bs,beam_size] descible length of sentence\n\n        normalized_scores= (new_scores/new_length.unsqueeze(-1)).view(self.batch_size,self.beam_size*V) #[bs,beam_size*V]\n        #normalized_scores = (new_scores).view(self.batch_size,self.beam_size * V)  # [bs,beam_size*V]\n        _,top_idx=torch.topk(normalized_scores,dim=-1,largest=True,k=self.beam_size)\n        beam_group=top_idx//V #[bs,beam_size]\n        token_idx=top_idx%V  #[bs,beam_size] #index of new tokens\n\n        #update scores, mask, length\n        self.sentence_length=beam_index_select(new_length,beam_group)\n        self.current_scores=torch.gather(new_scores.view(self.batch_size,self.beam_size*V),dim=1,index=top_idx)\n        self.finish_mask=torch.logical_or(beam_index_select(self.finish_mask,beam_group),token_idx==self.eos_idx)\n        self.current_sequence= beam_index_select(self.current_sequence.permute(1,2,0),beam_group).permute(2,0,1)\n\n        self.current_sequence=torch.cat([self.current_sequence,token_idx.unsqueeze(0)],dim=0)\n        self.current_cache=[self.rearrange_cache(beam_group,x) for x in new_cache]\n        return torch.all(self.finish_mask)\n```",
    "1336563": "Cangrats to be a GM!\nYour single models are so strong themselves! Can you give me a hint about progressive training in image size?\nI've tried same but didn't work due to positional encoding. How did you put multiple size pos enc to encoder output? (I also add pos enc to only k, v before attention)",
    "1336566": "Oh, the white space is not displayed well here. \nFor efficientnetv2, I train from 416 to 640 and then I finetune it on two sizes,  704 and 384x768.\nFor resnest101, I trained 416 and then finetune on two sizes, 640 and 384x768",
    "1336638": "But I just change to a new size and tune the model for 5 more epochs. The loss went down very quickly despite the encoding changed. \nPositional encodings on CNN output has small effects on the results, since CNN can learn positional information from zero padding.",
    "1336660": "thanks for sharing.  I used conda installation searched from online, some problem, plus other issues, I gave up.\n\nmy CV around 1.1 but LB 2.35.   Training without noise or augmentation?    leaders CV and LB are very close.",
    "1336667": "spent money on renting hardware resources. haha.",
    "1336880": "Thanks for the code",
    "1337191": "Congratulations on making it to GM! Can I ask - how did you get around the rdkit seg fault issue? I thought about using this in conjunction with beam search but I was worried that the code would just keep falling over.",
    "1337243": "I'm not the author, but I used multiprocessing Pool, that automatically restarts failed child processes:\n```\nfrom multiprocessing import Pool\nfrom multiprocessing.context import TimeoutError\n\ndef normalize_inchi(inchi):\n    try:\n        return Chem.MolToInchi(Chem.MolFromInchi(inchi))\n    except:\n        return None\n\npool = Pool(1)\nr = pool.apply_async(normalize_inchi, (inchi,))\ntry:\n    normalized_inchi = r.get(timeout=1)\nexcept TimeoutError:\n    normalized_inchi = None\n```",
    "1337404": "Hey, are you going to share the notebook for this cause I wanna check out the code.",
    "1337490": "stassl Ah yes, of course, clever! That's the way to do it. I had been doing the full test set using a bash script - checking the exit code and continuing to loop if it wasn't valid.",
    "1337689": "That's a nice solution, but in my case there is only 0-2 segmentation faults. So I just mannually rerun the code when it fails.",
    "1340781": "Congratulations."
  },
  "source": "meta"
}