{
  "id": 247661,
  "title": "27th place key highlights in 60 seconds",
  "url": "/competitions/bms-molecular-translation/discussion/247661",
  "author_name": "Alexander Soare",
  "post_date": "2021-06-20T15:26:23.723000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li>Architecture - Combination of:<ul>\n<li>VIT encoder with 8 transformer decoder layers. I felt on edge about having no inductive bias so I used 2x 3x3 convolutional kernels as input adaptors to the VIT (inconlusive whether it helped - doubt it). <strong>Edit</strong> Forgot to mention the important bit - the second one was stride = 2 thereby allowing me to do \"intelligent\" downsampling from larger input images than normal hardware constraints would permit.</li>\n<li>TNT encoder with 8 transformer decoder layers.</li></ul></li>\n<li>Used selective patches (thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the idea) making sure to only feed relevant patches to the encoder. This significantly sped up training.</li>\n<li>Salt and pepper (especially pepper) augmentation for the win! Check <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243932\" target=\"_blank\">this team's discussion</a>.</li>\n<li>RDKit generated images - I think this also helped because it let me increase training vocab.</li>\n<li><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalization script but with cascading and agreement:<ul>\n<li>Cascading = if the inchi was invalid, try the next best model's inchi. Huge improvement!</li>\n<li>Agreement = if top model inchi is different from 2nd, and 2nd == 3rd == 4th, then use the latter. Tiny improvement.</li></ul></li>\n<li>As <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> says - be patient and let the model train out for a while.</li>\n<li>Psuedolabelling training for 2 epochs - tiny improvement.</li>\n</ul>\n<p>What didn't work</p>\n<ul>\n<li>Kind of wasted my time a bit by splitting the transformer decoder up into two transformers, one for the chemical formula and another for the rest. The inspiration was similar to the one for <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243943\" target=\"_blank\">noise injection as this team did</a>. Too bad the idea didn't work.</li>\n</ul>\n<p>GG</p>",
  "messages": [
    {
      "id": 1358571,
      "postDate": "2021-06-20T15:26:23.723Z",
      "content": "<ul>\n<li>Architecture - Combination of:<ul>\n<li>VIT encoder with 8 transformer decoder layers. I felt on edge about having no inductive bias so I used 2x 3x3 convolutional kernels as input adaptors to the VIT (inconlusive whether it helped - doubt it). <strong>Edit</strong> Forgot to mention the important bit - the second one was stride = 2 thereby allowing me to do \"intelligent\" downsampling from larger input images than normal hardware constraints would permit.</li>\n<li>TNT encoder with 8 transformer decoder layers.</li></ul></li>\n<li>Used selective patches (thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the idea) making sure to only feed relevant patches to the encoder. This significantly sped up training.</li>\n<li>Salt and pepper (especially pepper) augmentation for the win! Check <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243932\" target=\"_blank\">this team's discussion</a>.</li>\n<li>RDKit generated images - I think this also helped because it let me increase training vocab.</li>\n<li><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalization script but with cascading and agreement:<ul>\n<li>Cascading = if the inchi was invalid, try the next best model's inchi. Huge improvement!</li>\n<li>Agreement = if top model inchi is different from 2nd, and 2nd == 3rd == 4th, then use the latter. Tiny improvement.</li></ul></li>\n<li>As <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> says - be patient and let the model train out for a while.</li>\n<li>Psuedolabelling training for 2 epochs - tiny improvement.</li>\n</ul>\n<p>What didn't work</p>\n<ul>\n<li>Kind of wasted my time a bit by splitting the transformer decoder up into two transformers, one for the chemical formula and another for the rest. The inspiration was similar to the one for <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243943\" target=\"_blank\">noise injection as this team did</a>. Too bad the idea didn't work.</li>\n</ul>\n<p>GG</p>",
      "rawMarkdown": "- Architecture - Combination of:\n  - VIT encoder with 8 transformer decoder layers. I felt on edge about having no inductive bias so I used 2x 3x3 convolutional kernels as input adaptors to the VIT (inconlusive whether it helped - doubt it). **Edit** Forgot to mention the important bit - the second one was stride = 2 thereby allowing me to do \"intelligent\" downsampling from larger input images than normal hardware constraints would permit.\n  - TNT encoder with 8 transformer decoder layers.\n- Used selective patches (thank you @hengck23 for the idea) making sure to only feed relevant patches to the encoder. This significantly sped up training.\n- Salt and pepper (especially pepper) augmentation for the win! Check [this team's discussion](https://www.kaggle.com/c/bms-molecular-translation/discussion/243932).\n- RDKit generated images - I think this also helped because it let me increase training vocab.\n- @nofreewill normalization script but with cascading and agreement:\n  - Cascading = if the inchi was invalid, try the next best model's inchi. Huge improvement!\n  - Agreement = if top model inchi is different from 2nd, and 2nd == 3rd == 4th, then use the latter. Tiny improvement.\n- As @fergusoci says - be patient and let the model train out for a while.\n- Psuedolabelling training for 2 epochs - tiny improvement.\n\nWhat didn't work\n- Kind of wasted my time a bit by splitting the transformer decoder up into two transformers, one for the chemical formula and another for the rest. The inspiration was similar to the one for [noise injection as this team did](https://www.kaggle.com/c/bms-molecular-translation/discussion/243943). Too bad the idea didn't work.\n\nGG\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1358571": "- Architecture - Combination of:\n  - VIT encoder with 8 transformer decoder layers. I felt on edge about having no inductive bias so I used 2x 3x3 convolutional kernels as input adaptors to the VIT (inconlusive whether it helped - doubt it). **Edit** Forgot to mention the important bit - the second one was stride = 2 thereby allowing me to do \"intelligent\" downsampling from larger input images than normal hardware constraints would permit.\n  - TNT encoder with 8 transformer decoder layers.\n- Used selective patches (thank you @hengck23 for the idea) making sure to only feed relevant patches to the encoder. This significantly sped up training.\n- Salt and pepper (especially pepper) augmentation for the win! Check [this team's discussion](https://www.kaggle.com/c/bms-molecular-translation/discussion/243932).\n- RDKit generated images - I think this also helped because it let me increase training vocab.\n- @nofreewill normalization script but with cascading and agreement:\n  - Cascading = if the inchi was invalid, try the next best model's inchi. Huge improvement!\n  - Agreement = if top model inchi is different from 2nd, and 2nd == 3rd == 4th, then use the latter. Tiny improvement.\n- As @fergusoci says - be patient and let the model train out for a while.\n- Psuedolabelling training for 2 epochs - tiny improvement.\n\nWhat didn't work\n- Kind of wasted my time a bit by splitting the transformer decoder up into two transformers, one for the chemical formula and another for the rest. The inspiration was similar to the one for [noise injection as this team did](https://www.kaggle.com/c/bms-molecular-translation/discussion/243943). Too bad the idea didn't work.\n\nGG\n"
  }
}