{
  "id": 223471,
  "title": "My idea to address this problem",
  "url": "/competitions/bms-molecular-translation/discussion/223471",
  "author_name": "",
  "post_date": "2021-03-04T02:25:40.824922700Z",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<p>After EDA, I found any InChI label can be create with a set of 38 tokens, those are:</p>\n<ol>\n<li>9 prefix: ['/B', '/C', '/b', '/c', '/h', '/i', '/m', '/s', '/t' ]</li>\n<li>29 unique characters: ['(', ')', '+', ',', '-', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', 'B', 'C', 'D', 'F', 'H', 'I', 'N', 'O', 'P', 'S', 'T', 'i', 'l', 'r']</li>\n</ol>\n<p>Maybe, we can use models like LSTM or Transformer to address this sequence generation problem.</p>\n<p>This just my intuition, i don't whether it works or not.</p>\n<p>-------------------------------Update ------------------------------------<br>\nDiscussion about sequence generation: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223419\" target=\"_blank\">Image Captioning - DL Research Papers + Code</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223212\" target=\"_blank\">End-to-end transformers approach with ViT and vanilla transformer decoder (idea)</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223293\" target=\"_blank\">Encoder - Decoder for Translation</a></li>\n<li>…</li>\n</ol>\n<p>If you want to know how to build a seq to seq generation model for image captioning, i recommmend:</p>\n<ol>\n<li><p><a href=\"https://github.com/yunjey/pytorch-tutorial/tree/master/tutorials/03-advanced/image_captioning\" target=\"_blank\">github: image captioning</a></p></li>\n<li><p><a href=\"https://suzyahyah.github.io/pytorch/2019/07/01/DataLoader-Pad-Pack-Sequence.html\" target=\"_blank\">Pad pack sequences for Pytorch batch processing with DataLoader</a>. Learn how to process various length of sequences with pytorch.</p></li>\n<li><p><a href=\"https://www.kaggle.com/kaushal2896/bms-mt-show-attend-and-tell-pytorch-baseline\" target=\"_blank\">BMS MT: Show, Attend, and Tell PyTorch Baseline</a>. Published notebook.</p></li>\n</ol>\n<p>Key steps to address this problem with basic Conv + LSTM model.</p>\n<ol>\n<li>build vocab to convert sequence into set of indexs.</li>\n<li>converting indexs into tensor with nn.Embedding</li>\n<li>build sequence generation model and training model with crossentropy loss. pytorch tool <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.utils.rnn.pack_padded_sequence.html?highlight=pack_padded_sequence#torch.nn.utils.rnn.pack_padded_sequence\" target=\"_blank\">pack padded sequence</a> will help you handle the various of input sequence, see link 2.</li>\n<li>during inference, we autoregressively produce predictions(a set of indexs) given an input image.</li>\n<li>converting the produced indexs back to text with vocab, then remove repeated prediction which is any character after the EOS symbol.</li>\n</ol>\n<p>I've published my implementation, see <a href=\"https://www.kaggle.com/wantsu/molecular-translation-ocsr-training\" target=\"_blank\">here</a><br>\nMy model use 10% training sample and score 75.1 on LB.</p>",
  "messages": [
    {
      "id": "1225856",
      "postDate": "03/04/2021 02:25:40",
      "content": "<p>After EDA, I found any InChI label can be create with a set of 38 tokens, those are:</p>\n<ol>\n<li>9 prefix: ['/B', '/C', '/b', '/c', '/h', '/i', '/m', '/s', '/t' ]</li>\n<li>29 unique characters: ['(', ')', '+', ',', '-', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', 'B', 'C', 'D', 'F', 'H', 'I', 'N', 'O', 'P', 'S', 'T', 'i', 'l', 'r']</li>\n</ol>\n<p>Maybe, we can use models like LSTM or Transformer to address this sequence generation problem.</p>\n<p>This just my intuition, i don't whether it works or not.</p>\n<p>-------------------------------Update ------------------------------------<br>\nDiscussion about sequence generation: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223419\" target=\"_blank\">Image Captioning - DL Research Papers + Code</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223212\" target=\"_blank\">End-to-end transformers approach with ViT and vanilla transformer decoder (idea)</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223293\" target=\"_blank\">Encoder - Decoder for Translation</a></li>\n<li>…</li>\n</ol>\n<p>If you want to know how to build a seq to seq generation model for image captioning, i recommmend:</p>\n<ol>\n<li><p><a href=\"https://github.com/yunjey/pytorch-tutorial/tree/master/tutorials/03-advanced/image_captioning\" target=\"_blank\">github: image captioning</a></p></li>\n<li><p><a href=\"https://suzyahyah.github.io/pytorch/2019/07/01/DataLoader-Pad-Pack-Sequence.html\" target=\"_blank\">Pad pack sequences for Pytorch batch processing with DataLoader</a>. Learn how to process various length of sequences with pytorch.</p></li>\n<li><p><a href=\"https://www.kaggle.com/kaushal2896/bms-mt-show-attend-and-tell-pytorch-baseline\" target=\"_blank\">BMS MT: Show, Attend, and Tell PyTorch Baseline</a>. Published notebook.</p></li>\n</ol>\n<p>Key steps to address this problem with basic Conv + LSTM model.</p>\n<ol>\n<li>build vocab to convert sequence into set of indexs.</li>\n<li>converting indexs into tensor with nn.Embedding</li>\n<li>build sequence generation model and training model with crossentropy loss. pytorch tool <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.utils.rnn.pack_padded_sequence.html?highlight=pack_padded_sequence#torch.nn.utils.rnn.pack_padded_sequence\" target=\"_blank\">pack padded sequence</a> will help you handle the various of input sequence, see link 2.</li>\n<li>during inference, we autoregressively produce predictions(a set of indexs) given an input image.</li>\n<li>converting the produced indexs back to text with vocab, then remove repeated prediction which is any character after the EOS symbol.</li>\n</ol>\n<p>I've published my implementation, see <a href=\"https://www.kaggle.com/wantsu/molecular-translation-ocsr-training\" target=\"_blank\">here</a><br>\nMy model use 10% training sample and score 75.1 on LB.</p>",
      "rawMarkdown": "After EDA, I found any InChI label can be create with a set of 38 tokens, those are:\n1. 9 prefix: ['/B', '/C', '/b', '/c', '/h', '/i', '/m', '/s', '/t' ]\n2. 29 unique characters: ['(', ')', '+', ',', '-', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', 'B', 'C', 'D', 'F', 'H', 'I', 'N', 'O', 'P', 'S', 'T', 'i', 'l', 'r']\n\nMaybe, we can use models like LSTM or Transformer to address this sequence generation problem.\n\nThis just my intuition, i don't whether it works or not.\n\n-------------------------------Update ------------------------------------\nDiscussion about sequence generation: \n1. [Image Captioning - DL Research Papers + Code](https://www.kaggle.com/c/bms-molecular-translation/discussion/223419)\n2. [End-to-end transformers approach with ViT and vanilla transformer decoder (idea)](https://www.kaggle.com/c/bms-molecular-translation/discussion/223212)\n3. [Encoder - Decoder for Translation](https://www.kaggle.com/c/bms-molecular-translation/discussion/223293)\n4. ...\n\nIf you want to know how to build a seq to seq generation model for image captioning, i recommmend:\n1. [github: image captioning](https://github.com/yunjey/pytorch-tutorial/tree/master/tutorials/03-advanced/image_captioning)\n2. [Pad pack sequences for Pytorch batch processing with DataLoader](https://suzyahyah.github.io/pytorch/2019/07/01/DataLoader-Pad-Pack-Sequence.html). Learn how to process various length of sequences with pytorch.\n\n3. [BMS MT: Show, Attend, and Tell PyTorch Baseline](https://www.kaggle.com/kaushal2896/bms-mt-show-attend-and-tell-pytorch-baseline). Published notebook.\n\nKey steps to address this problem with basic Conv + LSTM model.\n1. build vocab to convert sequence into set of indexs.\n2. converting indexs into tensor with nn.Embedding\n3. build sequence generation model and training model with crossentropy loss. pytorch tool [pack padded sequence](https://pytorch.org/docs/stable/generated/torch.nn.utils.rnn.pack_padded_sequence.html?highlight=pack_padded_sequence#torch.nn.utils.rnn.pack_padded_sequence) will help you handle the various of input sequence, see link 2.\n4. during inference, we autoregressively produce predictions(a set of indexs) given an input image.\n5. converting the produced indexs back to text with vocab, then remove repeated prediction which is any character after the EOS symbol.\n\nI've published my implementation, see [here] (https://www.kaggle.com/wantsu/molecular-translation-ocsr-training)\nMy model use 10% training sample and score 75.1 on LB.",
      "votes": null
    },
    {
      "id": "1226178",
      "postDate": "03/04/2021 10:02:50",
      "content": "<p>That seems like a logical idea to explore (esp. since people have used language models on chemical notation a lot, already), but I think it needs a good bit of tweaking. The other obvious way how to represent a molecule is a graph neural network, but I'm less familiar with using them for generation tasks, so a language model may well be an easier first try.</p>\n<p>I assume it will help to respect how the notation works  (see also <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223503\" target=\"_blank\">this forum post</a>):</p>\n<ul>\n<li>Let's look at a row in the training data: 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'</li>\n<li>The prefix \"InChI=1S/\" can basically be ignored since all target labels have it</li>\n<li>It seems like we almost need to generate three separate sequences and these need to be consistent: <ul>\n<li><strong>normal chemistry notation:</strong> \"C13H20OS\" and</li>\n<li><strong>\"connection layer\"</strong>: in what order the atoms (other than hydrogen atoms) are connected (and what the side branches are): \"c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14\" (you can basically omit the \"c\" and just look at the rest of this sequence) and</li>\n<li><strong>\"hydrogen layer\"</strong> how/how many hydrogen atoms are connected: \"h5-7,9,11,14H,8H2,1-4H3\" (you can basically omit the \"h\" and just look at the rest of this sequence)</li>\n<li>While we could of course train 3 separate language models for these, what we really need is something that ensures these are linked up and make sense (e.g. you cannot have more non-hydrogen atoms being connected in the connection layer than appear in the normal chemistry notation). </li>\n<li>Perhaps we can teach a transformer to pay attention to the right things to ensure this consistency (similarly, a LSTM probably would have long enough memory to deal with this, I suspect) or perhaps we can provide that feedback via the loss function by creating one that penalizes inconsistent behavior between the three different sequences that need to be generated (even if we generate them as one sequence). Or perhaps one could try an idea like triplet loss?!</li></ul></li>\n<li>I'm not sure the vocabulary you listed is the right one. E.g. \"Br\" is a single atom and probably each atom that gets mentioned should be its own \"word\" in the vocabulary.</li>\n<li>Another interesting how to use that \"-15-\" refers to 15 and that this is a number - rather than the letters \"1\" and \"5\" may be important - otherwise it'll be somewhat harder for a model to \"learn\" that 15 atoms are more than 2 and a lot less than 51 (especially, if 51 never appears in the training data). I think there's a decent amount of research suggesting that language models <a href=\"https://arxiv.org/pdf/1805.08154\" target=\"_blank\">struggle a bit with understanding numbers when they treat them like any other word in the vocabulary, but that we can improve on that</a>.</li>\n</ul>\n<p>The other big topic, of course, is the question is how you feed in the image. E.g. for a LSTM do you provide the image as an input at all time steps (which could be a bit computationally expensive)?</p>\n<p>Another difficulty: from the images, I'm not yet clear on how you know which atom is considered \"first\". As I <a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier\" target=\"_blank\">understand it, the string is actually unique</a> and presumably the <a href=\"https://depth-first.com/articles/2006/08/12/inchi-canonicalization-algorithm/\" target=\"_blank\">canonicalization</a> takes care of this? Making sure that one exactly understands that, will also be important, then one can perhaps post-process a model output to ensure it is exactly how it should be…</p>\n<p>This is certainly a nice challenging task!</p>",
      "rawMarkdown": "That seems like a logical idea to explore (esp. since people have used language models on chemical notation a lot, already), but I think it needs a good bit of tweaking. The other obvious way how to represent a molecule is a graph neural network, but I'm less familiar with using them for generation tasks, so a language model may well be an easier first try.\n\nI assume it will help to respect how the notation works  (see also [this forum post](https://www.kaggle.com/c/bms-molecular-translation/discussion/223503)):\n* Let's look at a row in the training data: 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\n* The prefix \"InChI=1S/\" can basically be ignored since all target labels have it\n* It seems like we almost need to generate three separate sequences and these need to be consistent: \n * **normal chemistry notation:** \"C13H20OS\" and\n * **\"connection layer\"**: in what order the atoms (other than hydrogen atoms) are connected (and what the side branches are): \"c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14\" (you can basically omit the \"c\" and just look at the rest of this sequence) and\n * **\"hydrogen layer\"** how/how many hydrogen atoms are connected: \"h5-7,9,11,14H,8H2,1-4H3\" (you can basically omit the \"h\" and just look at the rest of this sequence)\n * While we could of course train 3 separate language models for these, what we really need is something that ensures these are linked up and make sense (e.g. you cannot have more non-hydrogen atoms being connected in the connection layer than appear in the normal chemistry notation). \n * Perhaps we can teach a transformer to pay attention to the right things to ensure this consistency (similarly, a LSTM probably would have long enough memory to deal with this, I suspect) or perhaps we can provide that feedback via the loss function by creating one that penalizes inconsistent behavior between the three different sequences that need to be generated (even if we generate them as one sequence). Or perhaps one could try an idea like triplet loss?!\n* I'm not sure the vocabulary you listed is the right one. E.g. \"Br\" is a single atom and probably each atom that gets mentioned should be its own \"word\" in the vocabulary.\n* Another interesting how to use that \"-15-\" refers to 15 and that this is a number - rather than the letters \"1\" and \"5\" may be important - otherwise it'll be somewhat harder for a model to \"learn\" that 15 atoms are more than 2 and a lot less than 51 (especially, if 51 never appears in the training data). I think there's a decent amount of research suggesting that language models [struggle a bit with understanding numbers when they treat them like any other word in the vocabulary, but that we can improve on that](https://arxiv.org/pdf/1805.08154).\n\nThe other big topic, of course, is the question is how you feed in the image. E.g. for a LSTM do you provide the image as an input at all time steps (which could be a bit computationally expensive)?\n\nAnother difficulty: from the images, I'm not yet clear on how you know which atom is considered \"first\". As I [understand it, the string is actually unique](https://en.wikipedia.org/wiki/International_Chemical_Identifier) and presumably the [canonicalization](https://depth-first.com/articles/2006/08/12/inchi-canonicalization-algorithm/) takes care of this? Making sure that one exactly understands that, will also be important, then one can perhaps post-process a model output to ensure it is exactly how it should be...\n\nThis is certainly a nice challenging task!",
      "votes": null
    },
    {
      "id": "1226261",
      "postDate": "03/04/2021 11:58:56",
      "content": "<p>Great analysis!  <br>\nthese are someting I want to say:</p>\n<ol>\n<li>Your are right. 'Br‘ is a single atom, should be treat as a token. But i think autoregressive model will generate them together, that's is say if given last output 'B',  model will produce 'r'  at current time step since 'B' and 'r' are always together in corpus.</li>\n<li>I came across this idea because i think this problem is very similary to <strong>image captioning</strong>, so we can try to use sequence model and CNN to address this problem.<br>\n<img src=\"https://raw.githubusercontent.com/yunjey/pytorch-tutorial/master/tutorials/03-advanced/image_captioning/png/model.png\" alt=\"\"></li>\n<li>LSTM is weak to capture the long dependence and will suffer from high computation cost because of the long sequence. Transformer is the best choice.</li>\n</ol>\n<p>This is certainly a nice challenging task!</p>",
      "rawMarkdown": "Great analysis!  \nthese are someting I want to say:\n1. Your are right. 'Br‘ is a single atom, should be treat as a token. But i think autoregressive model will generate them together, that's is say if given last output 'B',  model will produce 'r'  at current time step since 'B' and 'r' are always together in corpus.\n2. I came across this idea because i think this problem is very similary to **image captioning**, so we can try to use sequence model and CNN to address this problem.\n![](https://raw.githubusercontent.com/yunjey/pytorch-tutorial/master/tutorials/03-advanced/image_captioning/png/model.png)\n3. LSTM is weak to capture the long dependence and will suffer from high computation cost because of the long sequence. Transformer is the best choice.\n\nThis is certainly a nice challenging task!",
      "votes": null
    },
    {
      "id": "1228009",
      "postDate": "03/06/2021 03:01:20",
      "content": "<p>I'm trying out this approach <a href=\"https://www.kaggle.com/arka47/convlstm-pytorch-skeleton\" target=\"_blank\">here</a> and need some implementation ideas, would be grateful if people come up and help:</p>\n<ol>\n<li><p>How exactly do I handle the string labels as tensors? Here I've tokenized them into a tensor of shape (batch_size, max_label_length, vocab_size), but then how exactly do I calculate the loss with the edit distance formula? Am I doing something wrong in the decoder part?</p></li>\n<li><p>Will ImageNet weights be any good for this problem? Else is there any other pre-training strategy that might be useful?</p></li>\n</ol>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "I'm trying out this approach [here](https://www.kaggle.com/arka47/convlstm-pytorch-skeleton) and need some implementation ideas, would be grateful if people come up and help:\n\n1. How exactly do I handle the string labels as tensors? Here I've tokenized them into a tensor of shape (batch_size, max_label_length, vocab_size), but then how exactly do I calculate the loss with the edit distance formula? Am I doing something wrong in the decoder part?\n\n2. Will ImageNet weights be any good for this problem? Else is there any other pre-training strategy that might be useful?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1228064",
      "postDate": "03/06/2021 04:52:54",
      "content": "<p>use crossentropy loss to train your model. I didn't seen any code to pack batch of sequence，how you handle different lengths among sequences？</p>",
      "rawMarkdown": "use crossentropy loss to train your model. I didn't seen any code to pack batch of sequence，how you handle different lengths among sequences？",
      "votes": null
    },
    {
      "id": "1229568",
      "postDate": "03/07/2021 12:42:52",
      "content": "<p>Let's say the length of the sequence is 100. Then after the 100th row, no row will contain a 1, all rows before that has exactly one 1 and the column indicates which character it is.</p>",
      "rawMarkdown": "Let's say the length of the sequence is 100. Then after the 100th row, no row will contain a 1, all rows before that has exactly one 1 and the column indicates which character it is.",
      "votes": null
    },
    {
      "id": "1232856",
      "postDate": "03/10/2021 02:59:49",
      "content": "<p><a href=\"https://www.kaggle.com/wantsu\" target=\"_blank\">@wantsu</a> Encoder - Decoder , Transformers , GANs and Image captioning models seems to be common suggestions coming from everyone in this forum . Thanks for sharing </p>",
      "rawMarkdown": "wantsu Encoder - Decoder , Transformers , GANs and Image captioning models seems to be common suggestions coming from everyone in this forum . Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1226178,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/04/2021 10:02:50",
      "content": "<p>That seems like a logical idea to explore (esp. since people have used language models on chemical notation a lot, already), but I think it needs a good bit of tweaking. The other obvious way how to represent a molecule is a graph neural network, but I'm less familiar with using them for generation tasks, so a language model may well be an easier first try.</p>\n<p>I assume it will help to respect how the notation works  (see also <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223503\" target=\"_blank\">this forum post</a>):</p>\n<ul>\n<li>Let's look at a row in the training data: 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'</li>\n<li>The prefix \"InChI=1S/\" can basically be ignored since all target labels have it</li>\n<li>It seems like we almost need to generate three separate sequences and these need to be consistent: <ul>\n<li><strong>normal chemistry notation:</strong> \"C13H20OS\" and</li>\n<li><strong>\"connection layer\"</strong>: in what order the atoms (other than hydrogen atoms) are connected (and what the side branches are): \"c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14\" (you can basically omit the \"c\" and just look at the rest of this sequence) and</li>\n<li><strong>\"hydrogen layer\"</strong> how/how many hydrogen atoms are connected: \"h5-7,9,11,14H,8H2,1-4H3\" (you can basically omit the \"h\" and just look at the rest of this sequence)</li>\n<li>While we could of course train 3 separate language models for these, what we really need is something that ensures these are linked up and make sense (e.g. you cannot have more non-hydrogen atoms being connected in the connection layer than appear in the normal chemistry notation). </li>\n<li>Perhaps we can teach a transformer to pay attention to the right things to ensure this consistency (similarly, a LSTM probably would have long enough memory to deal with this, I suspect) or perhaps we can provide that feedback via the loss function by creating one that penalizes inconsistent behavior between the three different sequences that need to be generated (even if we generate them as one sequence). Or perhaps one could try an idea like triplet loss?!</li></ul></li>\n<li>I'm not sure the vocabulary you listed is the right one. E.g. \"Br\" is a single atom and probably each atom that gets mentioned should be its own \"word\" in the vocabulary.</li>\n<li>Another interesting how to use that \"-15-\" refers to 15 and that this is a number - rather than the letters \"1\" and \"5\" may be important - otherwise it'll be somewhat harder for a model to \"learn\" that 15 atoms are more than 2 and a lot less than 51 (especially, if 51 never appears in the training data). I think there's a decent amount of research suggesting that language models <a href=\"https://arxiv.org/pdf/1805.08154\" target=\"_blank\">struggle a bit with understanding numbers when they treat them like any other word in the vocabulary, but that we can improve on that</a>.</li>\n</ul>\n<p>The other big topic, of course, is the question is how you feed in the image. E.g. for a LSTM do you provide the image as an input at all time steps (which could be a bit computationally expensive)?</p>\n<p>Another difficulty: from the images, I'm not yet clear on how you know which atom is considered \"first\". As I <a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier\" target=\"_blank\">understand it, the string is actually unique</a> and presumably the <a href=\"https://depth-first.com/articles/2006/08/12/inchi-canonicalization-algorithm/\" target=\"_blank\">canonicalization</a> takes care of this? Making sure that one exactly understands that, will also be important, then one can perhaps post-process a model output to ensure it is exactly how it should be…</p>\n<p>This is certainly a nice challenging task!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1226261,
          "author_name": "wantsu",
          "author_url": "",
          "post_date": "03/04/2021 11:58:56",
          "content": "<p>Great analysis!  <br>\nthese are someting I want to say:</p>\n<ol>\n<li>Your are right. 'Br‘ is a single atom, should be treat as a token. But i think autoregressive model will generate them together, that's is say if given last output 'B',  model will produce 'r'  at current time step since 'B' and 'r' are always together in corpus.</li>\n<li>I came across this idea because i think this problem is very similary to <strong>image captioning</strong>, so we can try to use sequence model and CNN to address this problem.<br>\n<img src=\"https://raw.githubusercontent.com/yunjey/pytorch-tutorial/master/tutorials/03-advanced/image_captioning/png/model.png\" alt=\"\"></li>\n<li>LSTM is weak to capture the long dependence and will suffer from high computation cost because of the long sequence. Transformer is the best choice.</li>\n</ol>\n<p>This is certainly a nice challenging task!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228009,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/06/2021 03:01:20",
          "content": "<p>I'm trying out this approach <a href=\"https://www.kaggle.com/arka47/convlstm-pytorch-skeleton\" target=\"_blank\">here</a> and need some implementation ideas, would be grateful if people come up and help:</p>\n<ol>\n<li><p>How exactly do I handle the string labels as tensors? Here I've tokenized them into a tensor of shape (batch_size, max_label_length, vocab_size), but then how exactly do I calculate the loss with the edit distance formula? Am I doing something wrong in the decoder part?</p></li>\n<li><p>Will ImageNet weights be any good for this problem? Else is there any other pre-training strategy that might be useful?</p></li>\n</ol>\n<p>Thanks in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228064,
          "author_name": "wantsu",
          "author_url": "",
          "post_date": "03/06/2021 04:52:54",
          "content": "<p>use crossentropy loss to train your model. I didn't seen any code to pack batch of sequence，how you handle different lengths among sequences？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1229568,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/07/2021 12:42:52",
          "content": "<p>Let's say the length of the sequence is 100. Then after the 100th row, no row will contain a 1, all rows before that has exactly one 1 and the column indicates which character it is.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1232856,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/10/2021 02:59:49",
      "content": "<p><a href=\"https://www.kaggle.com/wantsu\" target=\"_blank\">@wantsu</a> Encoder - Decoder , Transformers , GANs and Image captioning models seems to be common suggestions coming from everyone in this forum . Thanks for sharing </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1225856": "After EDA, I found any InChI label can be create with a set of 38 tokens, those are:\n1. 9 prefix: ['/B', '/C', '/b', '/c', '/h', '/i', '/m', '/s', '/t' ]\n2. 29 unique characters: ['(', ')', '+', ',', '-', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', 'B', 'C', 'D', 'F', 'H', 'I', 'N', 'O', 'P', 'S', 'T', 'i', 'l', 'r']\n\nMaybe, we can use models like LSTM or Transformer to address this sequence generation problem.\n\nThis just my intuition, i don't whether it works or not.\n\n-------------------------------Update ------------------------------------\nDiscussion about sequence generation: \n1. [Image Captioning - DL Research Papers + Code](https://www.kaggle.com/c/bms-molecular-translation/discussion/223419)\n2. [End-to-end transformers approach with ViT and vanilla transformer decoder (idea)](https://www.kaggle.com/c/bms-molecular-translation/discussion/223212)\n3. [Encoder - Decoder for Translation](https://www.kaggle.com/c/bms-molecular-translation/discussion/223293)\n4. ...\n\nIf you want to know how to build a seq to seq generation model for image captioning, i recommmend:\n1. [github: image captioning](https://github.com/yunjey/pytorch-tutorial/tree/master/tutorials/03-advanced/image_captioning)\n2. [Pad pack sequences for Pytorch batch processing with DataLoader](https://suzyahyah.github.io/pytorch/2019/07/01/DataLoader-Pad-Pack-Sequence.html). Learn how to process various length of sequences with pytorch.\n\n3. [BMS MT: Show, Attend, and Tell PyTorch Baseline](https://www.kaggle.com/kaushal2896/bms-mt-show-attend-and-tell-pytorch-baseline). Published notebook.\n\nKey steps to address this problem with basic Conv + LSTM model.\n1. build vocab to convert sequence into set of indexs.\n2. converting indexs into tensor with nn.Embedding\n3. build sequence generation model and training model with crossentropy loss. pytorch tool [pack padded sequence](https://pytorch.org/docs/stable/generated/torch.nn.utils.rnn.pack_padded_sequence.html?highlight=pack_padded_sequence#torch.nn.utils.rnn.pack_padded_sequence) will help you handle the various of input sequence, see link 2.\n4. during inference, we autoregressively produce predictions(a set of indexs) given an input image.\n5. converting the produced indexs back to text with vocab, then remove repeated prediction which is any character after the EOS symbol.\n\nI've published my implementation, see [here] (https://www.kaggle.com/wantsu/molecular-translation-ocsr-training)\nMy model use 10% training sample and score 75.1 on LB.",
    "1226178": "That seems like a logical idea to explore (esp. since people have used language models on chemical notation a lot, already), but I think it needs a good bit of tweaking. The other obvious way how to represent a molecule is a graph neural network, but I'm less familiar with using them for generation tasks, so a language model may well be an easier first try.\n\nI assume it will help to respect how the notation works  (see also [this forum post](https://www.kaggle.com/c/bms-molecular-translation/discussion/223503)):\n* Let's look at a row in the training data: 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\n* The prefix \"InChI=1S/\" can basically be ignored since all target labels have it\n* It seems like we almost need to generate three separate sequences and these need to be consistent: \n * **normal chemistry notation:** \"C13H20OS\" and\n * **\"connection layer\"**: in what order the atoms (other than hydrogen atoms) are connected (and what the side branches are): \"c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14\" (you can basically omit the \"c\" and just look at the rest of this sequence) and\n * **\"hydrogen layer\"** how/how many hydrogen atoms are connected: \"h5-7,9,11,14H,8H2,1-4H3\" (you can basically omit the \"h\" and just look at the rest of this sequence)\n * While we could of course train 3 separate language models for these, what we really need is something that ensures these are linked up and make sense (e.g. you cannot have more non-hydrogen atoms being connected in the connection layer than appear in the normal chemistry notation). \n * Perhaps we can teach a transformer to pay attention to the right things to ensure this consistency (similarly, a LSTM probably would have long enough memory to deal with this, I suspect) or perhaps we can provide that feedback via the loss function by creating one that penalizes inconsistent behavior between the three different sequences that need to be generated (even if we generate them as one sequence). Or perhaps one could try an idea like triplet loss?!\n* I'm not sure the vocabulary you listed is the right one. E.g. \"Br\" is a single atom and probably each atom that gets mentioned should be its own \"word\" in the vocabulary.\n* Another interesting how to use that \"-15-\" refers to 15 and that this is a number - rather than the letters \"1\" and \"5\" may be important - otherwise it'll be somewhat harder for a model to \"learn\" that 15 atoms are more than 2 and a lot less than 51 (especially, if 51 never appears in the training data). I think there's a decent amount of research suggesting that language models [struggle a bit with understanding numbers when they treat them like any other word in the vocabulary, but that we can improve on that](https://arxiv.org/pdf/1805.08154).\n\nThe other big topic, of course, is the question is how you feed in the image. E.g. for a LSTM do you provide the image as an input at all time steps (which could be a bit computationally expensive)?\n\nAnother difficulty: from the images, I'm not yet clear on how you know which atom is considered \"first\". As I [understand it, the string is actually unique](https://en.wikipedia.org/wiki/International_Chemical_Identifier) and presumably the [canonicalization](https://depth-first.com/articles/2006/08/12/inchi-canonicalization-algorithm/) takes care of this? Making sure that one exactly understands that, will also be important, then one can perhaps post-process a model output to ensure it is exactly how it should be...\n\nThis is certainly a nice challenging task!",
    "1226261": "Great analysis!  \nthese are someting I want to say:\n1. Your are right. 'Br‘ is a single atom, should be treat as a token. But i think autoregressive model will generate them together, that's is say if given last output 'B',  model will produce 'r'  at current time step since 'B' and 'r' are always together in corpus.\n2. I came across this idea because i think this problem is very similary to **image captioning**, so we can try to use sequence model and CNN to address this problem.\n![](https://raw.githubusercontent.com/yunjey/pytorch-tutorial/master/tutorials/03-advanced/image_captioning/png/model.png)\n3. LSTM is weak to capture the long dependence and will suffer from high computation cost because of the long sequence. Transformer is the best choice.\n\nThis is certainly a nice challenging task!",
    "1228009": "I'm trying out this approach [here](https://www.kaggle.com/arka47/convlstm-pytorch-skeleton) and need some implementation ideas, would be grateful if people come up and help:\n\n1. How exactly do I handle the string labels as tensors? Here I've tokenized them into a tensor of shape (batch_size, max_label_length, vocab_size), but then how exactly do I calculate the loss with the edit distance formula? Am I doing something wrong in the decoder part?\n\n2. Will ImageNet weights be any good for this problem? Else is there any other pre-training strategy that might be useful?\n\nThanks in advance!",
    "1228064": "use crossentropy loss to train your model. I didn't seen any code to pack batch of sequence，how you handle different lengths among sequences？",
    "1229568": "Let's say the length of the sequence is 100. Then after the 100th row, no row will contain a 1, all rows before that has exactly one 1 and the column indicates which character it is.",
    "1232856": "wantsu Encoder - Decoder , Transformers , GANs and Image captioning models seems to be common suggestions coming from everyone in this forum . Thanks for sharing"
  },
  "source": "meta"
}