{
  "id": 242354,
  "title": "Molecular Translation using Visual attention",
  "url": "/competitions/bms-molecular-translation/discussion/242354",
  "author_name": "ASHFAQUE",
  "post_date": "2021-05-28T14:55:50.183000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have created a python notebook for the molecular translation competition. I have used the visual attention model to generate the InChI name from the molecular structure image. This model has two parts, encoder, and decoder. For the encoder, I used a pre-trained Inceptionv3 model trained on the Imagenet dataset. For the decoder, I used the RNN model with GRU unit and attention mechanism.</p>\n<p>The same type of model is present on the Tensorflow website for Image Captioning: <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/image_captioning</a></p>\n<p>We can change the hyperparameters and the pre-trained model for encoding.</p>\n<p>I used the following parameters -</p>\n<p>BATCH_SIZE = 8<br>\nBUFFER_SIZE = 1000<br>\nembedding_dim = 256<br>\nunits = 512<br>\nvocab_size = 198<br>\nnum_steps = len(images) // BATCH_SIZE<br>\nfeatures_shape = 2048<br>\nattention_features_shape = 64</p>\n<p><strong>Tokenization of InChI names</strong></p>\n<p>The list of tokens that I used - </p>\n<p>TOKEN_LIST = [\"\", \"InChI=1S/\",\"\", \"\", \"/c\", \"/h\", \"/m\", \"/t\", \"/b\", \"/s\", \"/i\"] +\\<br>\n             ['Si', 'Br', 'Cl', 'F', 'I', 'N', 'O', 'P', 'S', 'C', 'H', 'B', ] +\\<br>\n             [str(i) for i in range(167,-1,-1)] +\\<br>\n             [\"+\", \"(\", \")\", \"-\", \",\", \"D\", \"T\"]</p>\n<p>For the tokenization, I used the re library.<br>\nFollowing is my function to convert into tokens</p>\n<p>def convert_to_tensor(label):<br>\n  token = [tok_2_int[\"\"]]<br>\n  l = label.split('/')<br>\n  token.append(tok_2_int[l[0]+'/'])<br>\n  f = re.split('(\\d+)', l[1])<br>\n  for c in f:<br>\n    if c.isnumeric()==False:<br>\n      st=0<br>\n      for i in range(len(c)+1):<br>\n        if c[st:i] in TOKEN_LIST:<br>\n          token.append(tok_2_int[c[st:i]])<br>\n          st=i<br>\n    else:<br>\n      token.append(tok_2_int[c])</p>\n<p>for i in range(2,len(l)):<br>\n    token.append(tok_2_int['/'+l[i][0]])<br>\n    s = re.split(r'(\\W+)', l[i][1:])<br>\n    for c in s:<br>\n      if c.isnumeric()==False and len(c)&gt;=2:<br>\n        if c[0] == '-' or c[0] == '+' or c[0] == ',' or c[0] == ')'  or c[0]== '(':<br>\n          for sp in c:<br>\n            token.append(tok_2_int[sp])<br>\n        else:<br>\n          cc = re.split('(\\d+)', c)<br>\n          for b in cc:<br>\n            if b.isnumeric()==False and len(b)&gt;=2:<br>\n              st=0<br>\n              for i in range(len(b)+1):<br>\n                if b[st:i] in TOKEN_LIST:<br>\n                  token.append(tok_2_int[b[st:i]])<br>\n                  st=i<br>\n            else:<br>\n               if len(b)&gt;0:<br>\n                 token.append(tok_2_int[b])<br>\n      else:<br>\n        if len(c)&gt;0:<br>\n          token.append(tok_2_int[c])<br>\n  token.append(tok_2_int[\"\"])<br>\n  return token</p>\n<p><strong>Training</strong></p>\n<p>During training, it took almost 30 min for 1 epoch using the batch of size 8 on Colab GPU.</p>",
  "messages": [
    {
      "id": 1326588,
      "postDate": "2021-05-28T14:55:50.183Z",
      "content": "<p>I have created a python notebook for the molecular translation competition. I have used the visual attention model to generate the InChI name from the molecular structure image. This model has two parts, encoder, and decoder. For the encoder, I used a pre-trained Inceptionv3 model trained on the Imagenet dataset. For the decoder, I used the RNN model with GRU unit and attention mechanism.</p>\n<p>The same type of model is present on the Tensorflow website for Image Captioning: <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/image_captioning</a></p>\n<p>We can change the hyperparameters and the pre-trained model for encoding.</p>\n<p>I used the following parameters -</p>\n<p>BATCH_SIZE = 8<br>\nBUFFER_SIZE = 1000<br>\nembedding_dim = 256<br>\nunits = 512<br>\nvocab_size = 198<br>\nnum_steps = len(images) // BATCH_SIZE<br>\nfeatures_shape = 2048<br>\nattention_features_shape = 64</p>\n<p><strong>Tokenization of InChI names</strong></p>\n<p>The list of tokens that I used - </p>\n<p>TOKEN_LIST = [\"\", \"InChI=1S/\",\"\", \"\", \"/c\", \"/h\", \"/m\", \"/t\", \"/b\", \"/s\", \"/i\"] +\\<br>\n             ['Si', 'Br', 'Cl', 'F', 'I', 'N', 'O', 'P', 'S', 'C', 'H', 'B', ] +\\<br>\n             [str(i) for i in range(167,-1,-1)] +\\<br>\n             [\"+\", \"(\", \")\", \"-\", \",\", \"D\", \"T\"]</p>\n<p>For the tokenization, I used the re library.<br>\nFollowing is my function to convert into tokens</p>\n<p>def convert_to_tensor(label):<br>\n  token = [tok_2_int[\"\"]]<br>\n  l = label.split('/')<br>\n  token.append(tok_2_int[l[0]+'/'])<br>\n  f = re.split('(\\d+)', l[1])<br>\n  for c in f:<br>\n    if c.isnumeric()==False:<br>\n      st=0<br>\n      for i in range(len(c)+1):<br>\n        if c[st:i] in TOKEN_LIST:<br>\n          token.append(tok_2_int[c[st:i]])<br>\n          st=i<br>\n    else:<br>\n      token.append(tok_2_int[c])</p>\n<p>for i in range(2,len(l)):<br>\n    token.append(tok_2_int['/'+l[i][0]])<br>\n    s = re.split(r'(\\W+)', l[i][1:])<br>\n    for c in s:<br>\n      if c.isnumeric()==False and len(c)&gt;=2:<br>\n        if c[0] == '-' or c[0] == '+' or c[0] == ',' or c[0] == ')'  or c[0]== '(':<br>\n          for sp in c:<br>\n            token.append(tok_2_int[sp])<br>\n        else:<br>\n          cc = re.split('(\\d+)', c)<br>\n          for b in cc:<br>\n            if b.isnumeric()==False and len(b)&gt;=2:<br>\n              st=0<br>\n              for i in range(len(b)+1):<br>\n                if b[st:i] in TOKEN_LIST:<br>\n                  token.append(tok_2_int[b[st:i]])<br>\n                  st=i<br>\n            else:<br>\n               if len(b)&gt;0:<br>\n                 token.append(tok_2_int[b])<br>\n      else:<br>\n        if len(c)&gt;0:<br>\n          token.append(tok_2_int[c])<br>\n  token.append(tok_2_int[\"\"])<br>\n  return token</p>\n<p><strong>Training</strong></p>\n<p>During training, it took almost 30 min for 1 epoch using the batch of size 8 on Colab GPU.</p>",
      "rawMarkdown": "I have created a python notebook for the molecular translation competition. I have used the visual attention model to generate the InChI name from the molecular structure image. This model has two parts, encoder, and decoder. For the encoder, I used a pre-trained Inceptionv3 model trained on the Imagenet dataset. For the decoder, I used the RNN model with GRU unit and attention mechanism.\n\nThe same type of model is present on the Tensorflow website for Image Captioning: https://www.tensorflow.org/tutorials/text/image_captioning\n\nWe can change the hyperparameters and the pre-trained model for encoding.\n\nI used the following parameters -\n\nBATCH_SIZE = 8\nBUFFER_SIZE = 1000\nembedding_dim = 256\nunits = 512\nvocab_size = 198\nnum_steps = len(images) // BATCH_SIZE\nfeatures_shape = 2048\nattention_features_shape = 64\n\n**Tokenization of InChI names**\n\nThe list of tokens that I used - \n\nTOKEN_LIST = [\"<PAD>\", \"InChI=1S/\",\"<START>\", \"<END>\", \"/c\", \"/h\", \"/m\", \"/t\", \"/b\", \"/s\", \"/i\"] +\\\n             ['Si', 'Br', 'Cl', 'F', 'I', 'N', 'O', 'P', 'S', 'C', 'H', 'B', ] +\\\n             [str(i) for i in range(167,-1,-1)] +\\\n             [\"+\", \"(\", \")\", \"-\", \",\", \"D\", \"T\"]\n\nFor the tokenization, I used the re library.\nFollowing is my function to convert into tokens\n\n def convert_to_tensor(label):\n  token = [tok_2_int[\"<START>\"]]\n  l = label.split('/')\n  token.append(tok_2_int[l[0]+'/'])\n  f = re.split('(\\d+)', l[1])\n  for c in f:\n    if c.isnumeric()==False:\n      st=0\n      for i in range(len(c)+1):\n        if c[st:i] in TOKEN_LIST:\n          token.append(tok_2_int[c[st:i]])\n          st=i\n    else:\n      token.append(tok_2_int[c])\n\n  for i in range(2,len(l)):\n    token.append(tok_2_int['/'+l[i][0]])\n    s = re.split(r'(\\W+)', l[i][1:])\n    for c in s:\n      if c.isnumeric()==False and len(c)>=2:\n        if c[0] == '-' or c[0] == '+' or c[0] == ',' or c[0] == ')'  or c[0]== '(':\n          for sp in c:\n            token.append(tok_2_int[sp])\n        else:\n          cc = re.split('(\\d+)', c)\n          for b in cc:\n            if b.isnumeric()==False and len(b)>=2:\n              st=0\n              for i in range(len(b)+1):\n                if b[st:i] in TOKEN_LIST:\n                  token.append(tok_2_int[b[st:i]])\n                  st=i\n            else:\n               if len(b)>0:\n                 token.append(tok_2_int[b])\n      else:\n        if len(c)>0:\n          token.append(tok_2_int[c])\n  token.append(tok_2_int[\"<END>\"])\n  return token\n\n\n**Training**\n\nDuring training, it took almost 30 min for 1 epoch using the batch of size 8 on Colab GPU.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1326588": "I have created a python notebook for the molecular translation competition. I have used the visual attention model to generate the InChI name from the molecular structure image. This model has two parts, encoder, and decoder. For the encoder, I used a pre-trained Inceptionv3 model trained on the Imagenet dataset. For the decoder, I used the RNN model with GRU unit and attention mechanism.\n\nThe same type of model is present on the Tensorflow website for Image Captioning: https://www.tensorflow.org/tutorials/text/image_captioning\n\nWe can change the hyperparameters and the pre-trained model for encoding.\n\nI used the following parameters -\n\nBATCH_SIZE = 8\nBUFFER_SIZE = 1000\nembedding_dim = 256\nunits = 512\nvocab_size = 198\nnum_steps = len(images) // BATCH_SIZE\nfeatures_shape = 2048\nattention_features_shape = 64\n\n**Tokenization of InChI names**\n\nThe list of tokens that I used - \n\nTOKEN_LIST = [\"<PAD>\", \"InChI=1S/\",\"<START>\", \"<END>\", \"/c\", \"/h\", \"/m\", \"/t\", \"/b\", \"/s\", \"/i\"] +\\\n             ['Si', 'Br', 'Cl', 'F', 'I', 'N', 'O', 'P', 'S', 'C', 'H', 'B', ] +\\\n             [str(i) for i in range(167,-1,-1)] +\\\n             [\"+\", \"(\", \")\", \"-\", \",\", \"D\", \"T\"]\n\nFor the tokenization, I used the re library.\nFollowing is my function to convert into tokens\n\n def convert_to_tensor(label):\n  token = [tok_2_int[\"<START>\"]]\n  l = label.split('/')\n  token.append(tok_2_int[l[0]+'/'])\n  f = re.split('(\\d+)', l[1])\n  for c in f:\n    if c.isnumeric()==False:\n      st=0\n      for i in range(len(c)+1):\n        if c[st:i] in TOKEN_LIST:\n          token.append(tok_2_int[c[st:i]])\n          st=i\n    else:\n      token.append(tok_2_int[c])\n\n  for i in range(2,len(l)):\n    token.append(tok_2_int['/'+l[i][0]])\n    s = re.split(r'(\\W+)', l[i][1:])\n    for c in s:\n      if c.isnumeric()==False and len(c)>=2:\n        if c[0] == '-' or c[0] == '+' or c[0] == ',' or c[0] == ')'  or c[0]== '(':\n          for sp in c:\n            token.append(tok_2_int[sp])\n        else:\n          cc = re.split('(\\d+)', c)\n          for b in cc:\n            if b.isnumeric()==False and len(b)>=2:\n              st=0\n              for i in range(len(b)+1):\n                if b[st:i] in TOKEN_LIST:\n                  token.append(tok_2_int[b[st:i]])\n                  st=i\n            else:\n               if len(b)>0:\n                 token.append(tok_2_int[b])\n      else:\n        if len(c)>0:\n          token.append(tok_2_int[c])\n  token.append(tok_2_int[\"<END>\"])\n  return token\n\n\n**Training**\n\nDuring training, it took almost 30 min for 1 epoch using the batch of size 8 on Colab GPU."
  }
}