{
  "id": 229720,
  "title": "some model suggestions",
  "url": "/competitions/bms-molecular-translation/discussion/229720",
  "author_name": "",
  "post_date": "2021-03-31T11:30:07.800455300Z",
  "votes": 43,
  "comment_count": 42,
  "views": 0,
  "content": "<p><img src=\"https://i.ibb.co/ZJgR3dQ/Selection-042.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1258135",
      "postDate": "03/31/2021 11:30:07",
      "content": "<p><img src=\"https://i.ibb.co/ZJgR3dQ/Selection-042.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/ZJgR3dQ/Selection-042.png)",
      "votes": null
    },
    {
      "id": "1258136",
      "postDate": "03/31/2021 11:30:22",
      "content": "<p><img src=\"https://i.ibb.co/CnBPyy3/Selection-041.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/CnBPyy3/Selection-041.png)",
      "votes": null
    },
    {
      "id": "1258137",
      "postDate": "03/31/2021 11:30:42",
      "content": "<p><img src=\"https://i.ibb.co/55Fs2Rd/Selection-043.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/55Fs2Rd/Selection-043.png)",
      "votes": null
    },
    {
      "id": "1258138",
      "postDate": "03/31/2021 11:31:04",
      "content": "<p><img src=\"https://i.ibb.co/WBHTbjZ/Selection-044.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/WBHTbjZ/Selection-044.png)",
      "votes": null
    },
    {
      "id": "1258171",
      "postDate": "03/31/2021 12:04:44",
      "content": "<p><img src=\"https://i.ibb.co/Dzz9ZMc/Selection-046.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/Dzz9ZMc/Selection-046.png)",
      "votes": null
    },
    {
      "id": "1258298",
      "postDate": "03/31/2021 14:13:56",
      "content": "<p>I've Implemented CNN-TCN with image attention. Comparing to CNN-LSTM it becomes terribly slow and unusable with my setup. I was thinking about transformer as decoder but I guess performance will be terrible.</p>",
      "rawMarkdown": "I've Implemented CNN-TCN with image attention. Comparing to CNN-LSTM it becomes terribly slow and unusable with my setup. I was thinking about transformer as decoder but I guess performance will be terrible.",
      "votes": null
    },
    {
      "id": "1258340",
      "postDate": "03/31/2021 14:44:03",
      "content": "<p><img src=\"https://i.ibb.co/mSdN0Gf/Selection-047.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/mSdN0Gf/Selection-047.png)",
      "votes": null
    },
    {
      "id": "1259037",
      "postDate": "04/01/2021 05:24:04",
      "content": "<p>Thanks for sharing such nice contents, May I ask for the tools you used for drawing these pictures? 🙏🙏</p>",
      "rawMarkdown": "Thanks for sharing such nice contents, May I ask for the tools you used for drawing these pictures? 🙏🙏",
      "votes": null
    },
    {
      "id": "1259167",
      "postDate": "04/01/2021 07:29:12",
      "content": "<p><a href=\"https://www.kaggle.com/xzy777\" target=\"_blank\">@xzy777</a>  They usually use Microsoft Power point to draw these architectures</p>",
      "rawMarkdown": "xzy777  They usually use Microsoft Power point to draw these architectures",
      "votes": null
    },
    {
      "id": "1259399",
      "postDate": "04/01/2021 11:17:23",
      "content": "<p><img src=\"https://i.ibb.co/J7KyXfV/Selection-063.png\" alt=\"\"></p>\n<p>the original CNN-attention-LSTM is slow and cannot be parallelised (e.g. for beam search)<br>\ni attempt to redesign it (this is concept only. i have not implemented it)</p>\n<p>the above shows stacked of 2 LSTM. One can stack up more layers. It is like iterative refinement using attention</p>",
      "rawMarkdown": "![](https://i.ibb.co/J7KyXfV/Selection-063.png)\n\nthe original CNN-attention-LSTM is slow and cannot be parallelised (e.g. for beam search)\ni attempt to redesign it (this is concept only. i have not implemented it)\n\n\nthe above shows stacked of 2 LSTM. One can stack up more layers. It is like iterative refinement using attention",
      "votes": null
    },
    {
      "id": "1259869",
      "postDate": "04/01/2021 17:56:15",
      "content": "<p>Good idea! Why trasformers are so powerful in computer vision task? Where do their benefits in quality come from?</p>",
      "rawMarkdown": "Good idea! Why trasformers are so powerful in computer vision task? Where do their benefits in quality come from?",
      "votes": null
    },
    {
      "id": "1262225",
      "postDate": "04/04/2021 01:47:53",
      "content": "<p><img src=\"https://d3i71xaburhd42.cloudfront.net/9fb5e3db385588f671b11cfc8bf18efb90ee7b19/4-Figure4-1.png\" alt=\"\"></p>\n<p>it turns out that there is a paper that uses 1d conv to replace RNN in image captioning<br>\n<a href=\"https://arxiv.org/pdf/1711.09151.pdf\" target=\"_blank\">https://arxiv.org/pdf/1711.09151.pdf</a><br>\nConvolutional Image Captioning</p>\n<p><a href=\"http://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/\" target=\"_blank\">http://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/</a></p>",
      "rawMarkdown": "![](https://d3i71xaburhd42.cloudfront.net/9fb5e3db385588f671b11cfc8bf18efb90ee7b19/4-Figure4-1.png)\n\nit turns out that there is a paper that uses 1d conv to replace RNN in image captioning\nhttps://arxiv.org/pdf/1711.09151.pdf\nConvolutional Image Captioning\n\nhttp://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/",
      "votes": null
    },
    {
      "id": "1262292",
      "postDate": "04/04/2021 04:36:42",
      "content": "<p>adventureous kagglers can try CLIP for pretraining</p>\n<p><a href=\"https://arxiv.org/pdf/2103.00020.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.00020.pdf</a><br>\n<a href=\"https://openai.com/blog/clip/\" target=\"_blank\">https://openai.com/blog/clip/</a></p>\n<p>Learning Transferable Visual Models From Natural Language Supervision<br>\n\" We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, \"</p>\n<p>\"CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. \"</p>",
      "rawMarkdown": "adventureous kagglers can try CLIP for pretraining\n\nhttps://arxiv.org/pdf/2103.00020.pdf\nhttps://openai.com/blog/clip/\n\nLearning Transferable Visual Models From Natural Language Supervision\n\" We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, \"\n\n\n\"CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. \"",
      "votes": null
    },
    {
      "id": "1264603",
      "postDate": "04/06/2021 09:46:28",
      "content": "<p>template for tf user<br>\n<a href=\"https://guillaumegenthial.github.io/image-to-latex.html\" target=\"_blank\">https://guillaumegenthial.github.io/image-to-latex.html</a><br>\n(check both part 1 and 2 of the web article)</p>\n<p><img src=\"https://i.ibb.co/stbvcQY/Selection-040.png\" alt=\"\"></p>",
      "rawMarkdown": "template for tf user\nhttps://guillaumegenthial.github.io/image-to-latex.html\n(check both part 1 and 2 of the web article)\n\n![](https://i.ibb.co/stbvcQY/Selection-040.png)",
      "votes": null
    },
    {
      "id": "1264858",
      "postDate": "04/06/2021 13:23:54",
      "content": "<p>I have tried this, it is quite difficult on the full dataset but seems like a really interesting way to go. I’m hoping I will have some time to work on it some more before this competition is over. </p>",
      "rawMarkdown": "I have tried this, it is quite difficult on the full dataset but seems like a really interesting way to go. I’m hoping I will have some time to work on it some more before this competition is over.",
      "votes": null
    },
    {
      "id": "1265226",
      "postDate": "04/06/2021 17:37:21",
      "content": "<p>i wonder if the followings make sense:</p>\n<ul>\n<li>train a normal seq predictor</li>\n<li>now modify the ground truth, e.g. reverse the sequence. train a model to predict the reverse sequence.</li>\n</ul>\n<p>if both the model gives the same results, then you know it is correct.</p>\n<p>there are several ways that you can modify the ground truth. based on the prediction of these several modification, can we get a code that can error correct itself?</p>\n<p>the idea is error correction of transmission codes in noisy channels, which is common in communication theory.</p>\n<p>e.g.<br>\n<a href=\"https://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/\" target=\"_blank\">https://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/</a></p>\n<p>eg  druchok_design_correction_control.pdf</p>\n<p>2.6 Error correction model<br>\nError correction in SMILES strings can be framed as a standard sequence-to-sequence learning problem,<br>\nwhich is traditionally solved with attention-based encoder-decoder models</p>\n<p>imagine if you use GAN or autoencoder to synthesis chemical formula, i wonder how do one ensure that the generated sample is of a valid molecule? </p>",
      "rawMarkdown": "i wonder if the followings make sense:\n- train a normal seq predictor\n- now modify the ground truth, e.g. reverse the sequence. train a model to predict the reverse sequence.\n\n\nif both the model gives the same results, then you know it is correct.\n\nthere are several ways that you can modify the ground truth. based on the prediction of these several modification, can we get a code that can error correct itself?\n\nthe idea is error correction of transmission codes in noisy channels, which is common in communication theory.\n\ne.g.\nhttps://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/\n\neg  druchok_design_correction_control.pdf\n\n2.6 Error correction model\nError correction in SMILES strings can be framed as a standard sequence-to-sequence learning problem,\nwhich is traditionally solved with attention-based encoder-decoder models\n\nimagine if you use GAN or autoencoder to synthesis chemical formula, i wonder how do one ensure that the generated sample is of a valid molecule?",
      "votes": null
    },
    {
      "id": "1265241",
      "postDate": "04/06/2021 17:51:33",
      "content": "<p>is an ensemble of several seq2seq models = another seq2se model</p>",
      "rawMarkdown": "is an ensemble of several seq2seq models = another seq2se model",
      "votes": null
    },
    {
      "id": "1265308",
      "postDate": "04/06/2021 19:28:00",
      "content": "<p>There was a mention of <a href=\"https://arxiv.org/abs/1805.11973\" target=\"_blank\">MolGAN</a> in another discussion thread (couldn't find it anymore).</p>\n<p>tldr: MolGAN was able to generate valid molecule with high probability. Maybe one can leverage the idea in MolGAN to regularize GAN or autoencoder to generate valid molecule.</p>",
      "rawMarkdown": "There was a mention of [MolGAN](https://arxiv.org/abs/1805.11973) in another discussion thread (couldn't find it anymore).\n\ntldr: MolGAN was able to generate valid molecule with high probability. Maybe one can leverage the idea in MolGAN to regularize GAN or autoencoder to generate valid molecule.",
      "votes": null
    },
    {
      "id": "1265489",
      "postDate": "04/06/2021 23:58:04",
      "content": "<p>an interesting way to decode:<br>\n<a href=\"https://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/\" target=\"_blank\">https://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/</a></p>\n<p>levenshtein-transformer<br>\n<img src=\"https://iconictranslation.com/wp-content/uploads/2020/05/NMT-82-LVT-fig-1.jpg\" alt=\"\"></p>",
      "rawMarkdown": "an interesting way to decode:\nhttps://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/\n\nlevenshtein-transformer\n![](https://iconictranslation.com/wp-content/uploads/2020/05/NMT-82-LVT-fig-1.jpg)",
      "votes": null
    },
    {
      "id": "1266690",
      "postDate": "04/08/2021 02:08:39",
      "content": "<p>test-time augmentation<br>\n<img src=\"https://i.ibb.co/tYnNsNV/Selection-050.png\" alt=\"\"></p>",
      "rawMarkdown": "test-time augmentation\n![](https://i.ibb.co/tYnNsNV/Selection-050.png)",
      "votes": null
    },
    {
      "id": "1269818",
      "postDate": "04/11/2021 00:07:15",
      "content": "<p><img src=\"https://miro.medium.com/max/875/0*l0J-2hNRpTBRRYQl\" alt=\"\"></p>\n<p>i almost forget there is this image bert that you can do multi-task training</p>\n<p>we now have a unique problem:</p>\n<ul>\n<li>we have image,text pair</li>\n<li>we have additional images without text (test data)</li>\n<li>we have additional text without images (external csv)</li>\n</ul>\n<p>how to use all data for pretraining?</p>",
      "rawMarkdown": "![](https://miro.medium.com/max/875/0*l0J-2hNRpTBRRYQl)\n\ni almost forget there is this image bert that you can do multi-task training\n\nwe now have a unique problem:\n- we have image,text pair\n- we have additional images without text (test data)\n- we have additional text without images (external csv)\n\nhow to use all data for pretraining?",
      "votes": null
    },
    {
      "id": "1269822",
      "postDate": "04/11/2021 00:20:48",
      "content": "<p>best of both worlds?<br>\nDeepMind, Microsoft, Allen AI &amp; UW Researchers Convert Pretrained Transformers into RNNs, Lowering Memory Cost While Retaining High Accuracy</p>\n<p><a href=\"https://arxiv.org/pdf/2103.13076.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.13076.pdf</a><br>\nwould use this if it is a code challenge </p>",
      "rawMarkdown": "best of both worlds?\nDeepMind, Microsoft, Allen AI & UW Researchers Convert Pretrained Transformers into RNNs, Lowering Memory Cost While Retaining High Accuracy\n\nhttps://arxiv.org/pdf/2103.13076.pdf\nwould use this if it is a code challenge",
      "votes": null
    },
    {
      "id": "1269823",
      "postDate": "04/11/2021 00:31:46",
      "content": "<p>MOCO-v3</p>\n<p>\"This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know<br>\nbaseline given the recent progress in computer vision: self-supervised learning for Visual Transformers (ViT).\"</p>\n<p>\"We also verify our models in GPUs using PyTorch. It takes 24 hours for ViT-B in 128 GPUs (vs. 2.1 hours in 256<br>\nTPUs).\"</p>\n<p>more here:<br>\n<a href=\"https://github.com/dk-liang/Awesome-Visual-Transformer\" target=\"_blank\">https://github.com/dk-liang/Awesome-Visual-Transformer</a></p>\n<hr>",
      "rawMarkdown": "MOCO-v3\n\n\"This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know\nbaseline given the recent progress in computer vision: self-supervised learning for Visual Transformers (ViT).\"\n\n\"We also verify our models in GPUs using PyTorch. It takes 24 hours for ViT-B in 128 GPUs (vs. 2.1 hours in 256\nTPUs).\"\n\nmore here:\nhttps://github.com/dk-liang/Awesome-Visual-Transformer\n\n---",
      "votes": null
    },
    {
      "id": "1269826",
      "postDate": "04/11/2021 00:41:31",
      "content": "<p>CPTR: FULL TRANSFORMER NETWORK FOR IMAGE CAPTIONING<br>\n<img src=\"https://pbs.twimg.com/media/EstJICeXEAMnt6y.jpg\" alt=\"\"></p>",
      "rawMarkdown": "CPTR: FULL TRANSFORMER NETWORK FOR IMAGE CAPTIONING\n![](https://pbs.twimg.com/media/EstJICeXEAMnt6y.jpg)",
      "votes": null
    },
    {
      "id": "1269829",
      "postDate": "04/11/2021 00:45:42",
      "content": "<p><img src=\"https://github.com/Sara-Ahmed/SiT/raw/main/SiT.png\" alt=\"\"></p>\n<p>SiT: Self-supervised image Transformer</p>",
      "rawMarkdown": "![](https://github.com/Sara-Ahmed/SiT/raw/main/SiT.png)\n\nSiT: Self-supervised image Transformer",
      "votes": null
    },
    {
      "id": "1269842",
      "postDate": "04/11/2021 01:07:54",
      "content": "<p>something like this</p>\n<p><img src=\"https://i.ibb.co/KFqdVkp/Selection-098.png\" alt=\"\"></p>",
      "rawMarkdown": "something like this\n\n![](https://i.ibb.co/KFqdVkp/Selection-098.png)",
      "votes": null
    },
    {
      "id": "1269844",
      "postDate": "04/11/2021 01:21:46",
      "content": "<p><img src=\"https://i.ibb.co/sF077j1/Selection-101.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/sF077j1/Selection-101.png)",
      "votes": null
    },
    {
      "id": "1270010",
      "postDate": "04/11/2021 06:59:24",
      "content": "<p><img src=\"https://i.ibb.co/KXpfd5z/Selection-104.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/KXpfd5z/Selection-104.png)",
      "votes": null
    },
    {
      "id": "1270016",
      "postDate": "04/11/2021 07:09:32",
      "content": "<p>is there a way to augment by combing 2 small molecules to a larger one? </p>",
      "rawMarkdown": "is there a way to augment by combing 2 small molecules to a larger one?",
      "votes": null
    },
    {
      "id": "1270090",
      "postDate": "04/11/2021 09:35:42",
      "content": "<p>I used an autoencoder to synthesise novel SMILES during my Masters dissertation. By sampling the latent space around known molecules it was possible to generate novel molecules. The closer the sampling to the known molecule, the more valid SMILES were produced, but at a lower uniqueness rate. </p>\n<p>Note that it is possible to write a chemical formula (e.g. SMILES or InChI) which is completely valid, but the molecule is unlikely to exist or be stable in the physical world. </p>",
      "rawMarkdown": "I used an autoencoder to synthesise novel SMILES during my Masters dissertation. By sampling the latent space around known molecules it was possible to generate novel molecules. The closer the sampling to the known molecule, the more valid SMILES were produced, but at a lower uniqueness rate. \n\nNote that it is possible to write a chemical formula (e.g. SMILES or InChI) which is completely valid, but the molecule is unlikely to exist or be stable in the physical world.",
      "votes": null
    },
    {
      "id": "1270283",
      "postDate": "04/11/2021 13:42:06",
      "content": "<p><img src=\"https://github.com/clovaai/SATRN/raw/master/figures/architecture.png\" alt=\"\"><br>\nOn Recognizing Texts of Arbitrary Shapes with 2D Self-Attention</p>\n<p>\"Adaptive 2D positional encoding (A2DPE) This new positional encoding is necessary for dynamically adapting to<br>\nthe inherent aspect ratios incurred by overall text alignment (horizontal, diagonal, or vertical). As alternative options, we consider not doing any positional encoding at all (“None”) (Zhang et al. 2019; Wang et al. 2017), using<br>\n1D positional encoding over flattened feature map (“1DFlatten”), using concatenation of height and width positional encodings (“2D-Concat”) (Parmar et al. 2018), and the A2DPE that we propose. See Table 3a for the results.<br>\nWe observe that A2DPE provides the best accuracy among four options considered\"</p>\n<p>this reminds me of the \"Spatial Transformer Networks\" which is used to align input using affine transformation.</p>\n<p>here we can learn a 2d positional encoding to make input image affine invariant</p>",
      "rawMarkdown": "![](https://github.com/clovaai/SATRN/raw/master/figures/architecture.png)\nOn Recognizing Texts of Arbitrary Shapes with 2D Self-Attention\n\n\"Adaptive 2D positional encoding (A2DPE) This new positional encoding is necessary for dynamically adapting to\nthe inherent aspect ratios incurred by overall text alignment (horizontal, diagonal, or vertical). As alternative options, we consider not doing any positional encoding at all (“None”) (Zhang et al. 2019; Wang et al. 2017), using\n1D positional encoding over flattened feature map (“1DFlatten”), using concatenation of height and width positional encodings (“2D-Concat”) (Parmar et al. 2018), and the A2DPE that we propose. See Table 3a for the results.\nWe observe that A2DPE provides the best accuracy among four options considered\"\n\n\nthis reminds me of the \"Spatial Transformer Networks\" which is used to align input using affine transformation.\n\nhere we can learn a 2d positional encoding to make input image affine invariant",
      "votes": null
    },
    {
      "id": "1270345",
      "postDate": "04/11/2021 14:51:15",
      "content": "<p>i realise that you actually don't have to input the image. you can just input the non-empy block.<br>\nhere you will have a variable length encoder.</p>",
      "rawMarkdown": "i realise that you actually don't have to input the image. you can just input the non-empy block.\nhere you will have a variable length encoder.",
      "votes": null
    },
    {
      "id": "1270358",
      "postDate": "04/11/2021 15:10:54",
      "content": "<p>Thanks for share! :)</p>",
      "rawMarkdown": "Thanks for share! :)",
      "votes": null
    },
    {
      "id": "1270366",
      "postDate": "04/11/2021 15:21:45",
      "content": "<p><img src=\"https://i.ibb.co/CQ980Hv/Selection-106.png\" alt=\"\"></p>\n<p>kinda of error correction?<br>\n(the input is noisy of course)</p>",
      "rawMarkdown": "![](https://i.ibb.co/CQ980Hv/Selection-106.png)\n\nkinda of error correction?\n(the input is noisy of course)",
      "votes": null
    },
    {
      "id": "1270714",
      "postDate": "04/11/2021 23:22:36",
      "content": "<p>good shares</p>",
      "rawMarkdown": "good shares",
      "votes": null
    },
    {
      "id": "1271925",
      "postDate": "04/13/2021 04:24:38",
      "content": "<p><img src=\"https://i.ibb.co/55DJqBc/Selection-029.png\" alt=\"\"><br>\ntransformer graph decoder</p>\n<p><a href=\"https://arxiv.org/pdf/2006.05213.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.05213.pdf</a> </p>",
      "rawMarkdown": "![](https://i.ibb.co/55DJqBc/Selection-029.png)\ntransformer graph decoder\n\nhttps://arxiv.org/pdf/2006.05213.pdf",
      "votes": null
    },
    {
      "id": "1273003",
      "postDate": "04/14/2021 02:09:34",
      "content": "<p>graph-based encoder</p>\n<p><img src=\"https://i.ibb.co/m5c9TBJ/Selection-032.png\" alt=\"\"></p>",
      "rawMarkdown": "graph-based encoder\n\n![](https://i.ibb.co/m5c9TBJ/Selection-032.png)",
      "votes": null
    },
    {
      "id": "1273456",
      "postDate": "04/14/2021 11:16:21",
      "content": "<p><a href=\"https://molvec.ncats.io/\" target=\"_blank\">https://molvec.ncats.io/</a><br>\n<a href=\"https://github.com/ncats/molvec-web\" target=\"_blank\">https://github.com/ncats/molvec-web</a></p>\n<p>this should be combined with a transformer</p>\n<p><img src=\"https://i.ibb.co/x6j7Hjp/Selection-031.png\" alt=\"\"></p>",
      "rawMarkdown": "https://molvec.ncats.io/\nhttps://github.com/ncats/molvec-web\n\n\nthis should be combined with a transformer\n\n![](https://i.ibb.co/x6j7Hjp/Selection-031.png)",
      "votes": null
    },
    {
      "id": "1275982",
      "postDate": "04/17/2021 00:18:32",
      "content": "<p><a href=\"https://davide-belli.github.io/generative-graph-transformer.html\" target=\"_blank\">https://davide-belli.github.io/generative-graph-transformer.html</a><br>\n<img src=\"https://davide-belli.github.io/images/blog-ggt/architecture.png\" alt=\"\"></p>",
      "rawMarkdown": "https://davide-belli.github.io/generative-graph-transformer.html\n![](https://davide-belli.github.io/images/blog-ggt/architecture.png)",
      "votes": null
    },
    {
      "id": "1292458",
      "postDate": "05/04/2021 01:24:47",
      "content": "<p>going multi-task<br>\n<img src=\"https://i.ibb.co/qyBj8cL/Selection-092.png\" alt=\"\"></p>",
      "rawMarkdown": "going multi-task\n![](https://i.ibb.co/qyBj8cL/Selection-092.png)",
      "votes": null
    },
    {
      "id": "1292953",
      "postDate": "05/04/2021 12:46:23",
      "content": "<p>multi-task learning is probably <br>\n\"The Feynman Technique\"<br>\n<a href=\"https://www.youtube.com/watch?v=6QUjy0B3lEw\" target=\"_blank\">https://www.youtube.com/watch?v=6QUjy0B3lEw</a></p>\n<p>为什么费曼技巧被称为终极学习法 (Why feynman technique works)<br>\n<a href=\"https://www.youtube.com/watch?v=7iNJyEbYDdc\" target=\"_blank\">https://www.youtube.com/watch?v=7iNJyEbYDdc</a></p>\n<p>\"let the machine teach the human and not human teaches machine\"<br>\nknowledge distillation should benefit the teacher and not the student …</p>\n<hr>\n<p>in multi-task we care about input encoding. we care less about the task.</p>\n<p>we are forced to seek simple features (information compression) that can solve many tasks …</p>",
      "rawMarkdown": "multi-task learning is probably \n\"The Feynman Technique\"\nhttps://www.youtube.com/watch?v=6QUjy0B3lEw\n\n为什么费曼技巧被称为终极学习法 (Why feynman technique works)\nhttps://www.youtube.com/watch?v=7iNJyEbYDdc\n\n\"let the machine teach the human and not human teaches machine\"\nknowledge distillation should benefit the teacher and not the student ...\n\n---\n\nin multi-task we care about input encoding. we care less about the task.\n\nwe are forced to seek simple features (information compression) that can solve many tasks ...",
      "votes": null
    },
    {
      "id": "1295834",
      "postDate": "05/06/2021 18:19:54",
      "content": "<p>How do think MLP-mixer for this application? Can it replace Transfomer? Thanks.<br>\n<a href=\"https://arxiv.org/pdf/2105.01601.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.01601.pdf</a></p>",
      "rawMarkdown": "How do think MLP-mixer for this application? Can it replace Transfomer? Thanks.\nhttps://arxiv.org/pdf/2105.01601.pdf",
      "votes": null
    },
    {
      "id": "1303204",
      "postDate": "05/12/2021 00:46:46",
      "content": "<p><a href=\"https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/\" target=\"_blank\">https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/</a></p>\n<p><img src=\"https://blogs.biomedcentral.com/on-physicalsciences/wp-content/uploads/sites/14/2020/02/JCheminf_figure_1_rnn_sampling.gif\" alt=\"\"></p>",
      "rawMarkdown": "https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/\n\n\n![](https://blogs.biomedcentral.com/on-physicalsciences/wp-content/uploads/sites/14/2020/02/JCheminf_figure_1_rnn_sampling.gif)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1258136,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/31/2021 11:30:22",
      "content": "<p><img src=\"https://i.ibb.co/CnBPyy3/Selection-041.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1258137,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/31/2021 11:30:42",
      "content": "<p><img src=\"https://i.ibb.co/55Fs2Rd/Selection-043.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1258340,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/31/2021 14:44:03",
          "content": "<p><img src=\"https://i.ibb.co/mSdN0Gf/Selection-047.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1258138,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/31/2021 11:31:04",
      "content": "<p><img src=\"https://i.ibb.co/WBHTbjZ/Selection-044.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1258171,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/31/2021 12:04:44",
      "content": "<p><img src=\"https://i.ibb.co/Dzz9ZMc/Selection-046.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1258298,
      "author_name": "ubique",
      "author_url": "",
      "post_date": "03/31/2021 14:13:56",
      "content": "<p>I've Implemented CNN-TCN with image attention. Comparing to CNN-LSTM it becomes terribly slow and unusable with my setup. I was thinking about transformer as decoder but I guess performance will be terrible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1259037,
      "author_name": "xzy777",
      "author_url": "",
      "post_date": "04/01/2021 05:24:04",
      "content": "<p>Thanks for sharing such nice contents, May I ask for the tools you used for drawing these pictures? 🙏🙏</p>",
      "votes": null,
      "replies": [
        {
          "id": 1259167,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "04/01/2021 07:29:12",
          "content": "<p><a href=\"https://www.kaggle.com/xzy777\" target=\"_blank\">@xzy777</a>  They usually use Microsoft Power point to draw these architectures</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1259399,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/01/2021 11:17:23",
      "content": "<p><img src=\"https://i.ibb.co/J7KyXfV/Selection-063.png\" alt=\"\"></p>\n<p>the original CNN-attention-LSTM is slow and cannot be parallelised (e.g. for beam search)<br>\ni attempt to redesign it (this is concept only. i have not implemented it)</p>\n<p>the above shows stacked of 2 LSTM. One can stack up more layers. It is like iterative refinement using attention</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1259869,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "04/01/2021 17:56:15",
      "content": "<p>Good idea! Why trasformers are so powerful in computer vision task? Where do their benefits in quality come from?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1262225,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/04/2021 01:47:53",
      "content": "<p><img src=\"https://d3i71xaburhd42.cloudfront.net/9fb5e3db385588f671b11cfc8bf18efb90ee7b19/4-Figure4-1.png\" alt=\"\"></p>\n<p>it turns out that there is a paper that uses 1d conv to replace RNN in image captioning<br>\n<a href=\"https://arxiv.org/pdf/1711.09151.pdf\" target=\"_blank\">https://arxiv.org/pdf/1711.09151.pdf</a><br>\nConvolutional Image Captioning</p>\n<p><a href=\"http://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/\" target=\"_blank\">http://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1262292,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/04/2021 04:36:42",
      "content": "<p>adventureous kagglers can try CLIP for pretraining</p>\n<p><a href=\"https://arxiv.org/pdf/2103.00020.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.00020.pdf</a><br>\n<a href=\"https://openai.com/blog/clip/\" target=\"_blank\">https://openai.com/blog/clip/</a></p>\n<p>Learning Transferable Visual Models From Natural Language Supervision<br>\n\" We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, \"</p>\n<p>\"CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. \"</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264858,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/06/2021 13:23:54",
          "content": "<p>I have tried this, it is quite difficult on the full dataset but seems like a really interesting way to go. I’m hoping I will have some time to work on it some more before this competition is over. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264603,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/06/2021 09:46:28",
      "content": "<p>template for tf user<br>\n<a href=\"https://guillaumegenthial.github.io/image-to-latex.html\" target=\"_blank\">https://guillaumegenthial.github.io/image-to-latex.html</a><br>\n(check both part 1 and 2 of the web article)</p>\n<p><img src=\"https://i.ibb.co/stbvcQY/Selection-040.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1265226,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/06/2021 17:37:21",
      "content": "<p>i wonder if the followings make sense:</p>\n<ul>\n<li>train a normal seq predictor</li>\n<li>now modify the ground truth, e.g. reverse the sequence. train a model to predict the reverse sequence.</li>\n</ul>\n<p>if both the model gives the same results, then you know it is correct.</p>\n<p>there are several ways that you can modify the ground truth. based on the prediction of these several modification, can we get a code that can error correct itself?</p>\n<p>the idea is error correction of transmission codes in noisy channels, which is common in communication theory.</p>\n<p>e.g.<br>\n<a href=\"https://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/\" target=\"_blank\">https://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/</a></p>\n<p>eg  druchok_design_correction_control.pdf</p>\n<p>2.6 Error correction model<br>\nError correction in SMILES strings can be framed as a standard sequence-to-sequence learning problem,<br>\nwhich is traditionally solved with attention-based encoder-decoder models</p>\n<p>imagine if you use GAN or autoencoder to synthesis chemical formula, i wonder how do one ensure that the generated sample is of a valid molecule? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1265241,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/06/2021 17:51:33",
          "content": "<p>is an ensemble of several seq2seq models = another seq2se model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1265308,
          "author_name": "jy2tong",
          "author_url": "",
          "post_date": "04/06/2021 19:28:00",
          "content": "<p>There was a mention of <a href=\"https://arxiv.org/abs/1805.11973\" target=\"_blank\">MolGAN</a> in another discussion thread (couldn't find it anymore).</p>\n<p>tldr: MolGAN was able to generate valid molecule with high probability. Maybe one can leverage the idea in MolGAN to regularize GAN or autoencoder to generate valid molecule.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270090,
          "author_name": "talktocharles",
          "author_url": "",
          "post_date": "04/11/2021 09:35:42",
          "content": "<p>I used an autoencoder to synthesise novel SMILES during my Masters dissertation. By sampling the latent space around known molecules it was possible to generate novel molecules. The closer the sampling to the known molecule, the more valid SMILES were produced, but at a lower uniqueness rate. </p>\n<p>Note that it is possible to write a chemical formula (e.g. SMILES or InChI) which is completely valid, but the molecule is unlikely to exist or be stable in the physical world. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265489,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/06/2021 23:58:04",
      "content": "<p>an interesting way to decode:<br>\n<a href=\"https://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/\" target=\"_blank\">https://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/</a></p>\n<p>levenshtein-transformer<br>\n<img src=\"https://iconictranslation.com/wp-content/uploads/2020/05/NMT-82-LVT-fig-1.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1266690,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/08/2021 02:08:39",
      "content": "<p>test-time augmentation<br>\n<img src=\"https://i.ibb.co/tYnNsNV/Selection-050.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1269818,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 00:07:15",
      "content": "<p><img src=\"https://miro.medium.com/max/875/0*l0J-2hNRpTBRRYQl\" alt=\"\"></p>\n<p>i almost forget there is this image bert that you can do multi-task training</p>\n<p>we now have a unique problem:</p>\n<ul>\n<li>we have image,text pair</li>\n<li>we have additional images without text (test data)</li>\n<li>we have additional text without images (external csv)</li>\n</ul>\n<p>how to use all data for pretraining?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1269842,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/11/2021 01:07:54",
          "content": "<p>something like this</p>\n<p><img src=\"https://i.ibb.co/KFqdVkp/Selection-098.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1269822,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 00:20:48",
      "content": "<p>best of both worlds?<br>\nDeepMind, Microsoft, Allen AI &amp; UW Researchers Convert Pretrained Transformers into RNNs, Lowering Memory Cost While Retaining High Accuracy</p>\n<p><a href=\"https://arxiv.org/pdf/2103.13076.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.13076.pdf</a><br>\nwould use this if it is a code challenge </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1269823,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 00:31:46",
      "content": "<p>MOCO-v3</p>\n<p>\"This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know<br>\nbaseline given the recent progress in computer vision: self-supervised learning for Visual Transformers (ViT).\"</p>\n<p>\"We also verify our models in GPUs using PyTorch. It takes 24 hours for ViT-B in 128 GPUs (vs. 2.1 hours in 256<br>\nTPUs).\"</p>\n<p>more here:<br>\n<a href=\"https://github.com/dk-liang/Awesome-Visual-Transformer\" target=\"_blank\">https://github.com/dk-liang/Awesome-Visual-Transformer</a></p>\n<hr>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1269826,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 00:41:31",
      "content": "<p>CPTR: FULL TRANSFORMER NETWORK FOR IMAGE CAPTIONING<br>\n<img src=\"https://pbs.twimg.com/media/EstJICeXEAMnt6y.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1269829,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 00:45:42",
      "content": "<p><img src=\"https://github.com/Sara-Ahmed/SiT/raw/main/SiT.png\" alt=\"\"></p>\n<p>SiT: Self-supervised image Transformer</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1269844,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 01:21:46",
      "content": "<p><img src=\"https://i.ibb.co/sF077j1/Selection-101.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1270010,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/11/2021 06:59:24",
          "content": "<p><img src=\"https://i.ibb.co/KXpfd5z/Selection-104.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270345,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/11/2021 14:51:15",
          "content": "<p>i realise that you actually don't have to input the image. you can just input the non-empy block.<br>\nhere you will have a variable length encoder.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1270016,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 07:09:32",
      "content": "<p>is there a way to augment by combing 2 small molecules to a larger one? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1270283,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 13:42:06",
      "content": "<p><img src=\"https://github.com/clovaai/SATRN/raw/master/figures/architecture.png\" alt=\"\"><br>\nOn Recognizing Texts of Arbitrary Shapes with 2D Self-Attention</p>\n<p>\"Adaptive 2D positional encoding (A2DPE) This new positional encoding is necessary for dynamically adapting to<br>\nthe inherent aspect ratios incurred by overall text alignment (horizontal, diagonal, or vertical). As alternative options, we consider not doing any positional encoding at all (“None”) (Zhang et al. 2019; Wang et al. 2017), using<br>\n1D positional encoding over flattened feature map (“1DFlatten”), using concatenation of height and width positional encodings (“2D-Concat”) (Parmar et al. 2018), and the A2DPE that we propose. See Table 3a for the results.<br>\nWe observe that A2DPE provides the best accuracy among four options considered\"</p>\n<p>this reminds me of the \"Spatial Transformer Networks\" which is used to align input using affine transformation.</p>\n<p>here we can learn a 2d positional encoding to make input image affine invariant</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1270358,
      "author_name": "dasmehdixtr",
      "author_url": "",
      "post_date": "04/11/2021 15:10:54",
      "content": "<p>Thanks for share! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1270366,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 15:21:45",
      "content": "<p><img src=\"https://i.ibb.co/CQ980Hv/Selection-106.png\" alt=\"\"></p>\n<p>kinda of error correction?<br>\n(the input is noisy of course)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1270714,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "04/11/2021 23:22:36",
      "content": "<p>good shares</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1271925,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/13/2021 04:24:38",
      "content": "<p><img src=\"https://i.ibb.co/55DJqBc/Selection-029.png\" alt=\"\"><br>\ntransformer graph decoder</p>\n<p><a href=\"https://arxiv.org/pdf/2006.05213.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.05213.pdf</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1273003,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/14/2021 02:09:34",
      "content": "<p>graph-based encoder</p>\n<p><img src=\"https://i.ibb.co/m5c9TBJ/Selection-032.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1273456,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/14/2021 11:16:21",
      "content": "<p><a href=\"https://molvec.ncats.io/\" target=\"_blank\">https://molvec.ncats.io/</a><br>\n<a href=\"https://github.com/ncats/molvec-web\" target=\"_blank\">https://github.com/ncats/molvec-web</a></p>\n<p>this should be combined with a transformer</p>\n<p><img src=\"https://i.ibb.co/x6j7Hjp/Selection-031.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1275982,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/17/2021 00:18:32",
      "content": "<p><a href=\"https://davide-belli.github.io/generative-graph-transformer.html\" target=\"_blank\">https://davide-belli.github.io/generative-graph-transformer.html</a><br>\n<img src=\"https://davide-belli.github.io/images/blog-ggt/architecture.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1292458,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/04/2021 01:24:47",
      "content": "<p>going multi-task<br>\n<img src=\"https://i.ibb.co/qyBj8cL/Selection-092.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1292953,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/04/2021 12:46:23",
          "content": "<p>multi-task learning is probably <br>\n\"The Feynman Technique\"<br>\n<a href=\"https://www.youtube.com/watch?v=6QUjy0B3lEw\" target=\"_blank\">https://www.youtube.com/watch?v=6QUjy0B3lEw</a></p>\n<p>为什么费曼技巧被称为终极学习法 (Why feynman technique works)<br>\n<a href=\"https://www.youtube.com/watch?v=7iNJyEbYDdc\" target=\"_blank\">https://www.youtube.com/watch?v=7iNJyEbYDdc</a></p>\n<p>\"let the machine teach the human and not human teaches machine\"<br>\nknowledge distillation should benefit the teacher and not the student …</p>\n<hr>\n<p>in multi-task we care about input encoding. we care less about the task.</p>\n<p>we are forced to seek simple features (information compression) that can solve many tasks …</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1295834,
      "author_name": "ybwu01",
      "author_url": "",
      "post_date": "05/06/2021 18:19:54",
      "content": "<p>How do think MLP-mixer for this application? Can it replace Transfomer? Thanks.<br>\n<a href=\"https://arxiv.org/pdf/2105.01601.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.01601.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1303204,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/12/2021 00:46:46",
      "content": "<p><a href=\"https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/\" target=\"_blank\">https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/</a></p>\n<p><img src=\"https://blogs.biomedcentral.com/on-physicalsciences/wp-content/uploads/sites/14/2020/02/JCheminf_figure_1_rnn_sampling.gif\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258135": "![](https://i.ibb.co/ZJgR3dQ/Selection-042.png)",
    "1258136": "![](https://i.ibb.co/CnBPyy3/Selection-041.png)",
    "1258137": "![](https://i.ibb.co/55Fs2Rd/Selection-043.png)",
    "1258138": "![](https://i.ibb.co/WBHTbjZ/Selection-044.png)",
    "1258171": "![](https://i.ibb.co/Dzz9ZMc/Selection-046.png)",
    "1258298": "I've Implemented CNN-TCN with image attention. Comparing to CNN-LSTM it becomes terribly slow and unusable with my setup. I was thinking about transformer as decoder but I guess performance will be terrible.",
    "1258340": "![](https://i.ibb.co/mSdN0Gf/Selection-047.png)",
    "1259037": "Thanks for sharing such nice contents, May I ask for the tools you used for drawing these pictures? 🙏🙏",
    "1259167": "xzy777  They usually use Microsoft Power point to draw these architectures",
    "1259399": "![](https://i.ibb.co/J7KyXfV/Selection-063.png)\n\nthe original CNN-attention-LSTM is slow and cannot be parallelised (e.g. for beam search)\ni attempt to redesign it (this is concept only. i have not implemented it)\n\n\nthe above shows stacked of 2 LSTM. One can stack up more layers. It is like iterative refinement using attention",
    "1259869": "Good idea! Why trasformers are so powerful in computer vision task? Where do their benefits in quality come from?",
    "1262225": "![](https://d3i71xaburhd42.cloudfront.net/9fb5e3db385588f671b11cfc8bf18efb90ee7b19/4-Figure4-1.png)\n\nit turns out that there is a paper that uses 1d conv to replace RNN in image captioning\nhttps://arxiv.org/pdf/1711.09151.pdf\nConvolutional Image Captioning\n\nhttp://visal.cs.cityu.edu.hk/research/convolutional-decoders-for-image-captioning/",
    "1262292": "adventureous kagglers can try CLIP for pretraining\n\nhttps://arxiv.org/pdf/2103.00020.pdf\nhttps://openai.com/blog/clip/\n\nLearning Transferable Visual Models From Natural Language Supervision\n\" We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, \"\n\n\n\"CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. \"",
    "1264603": "template for tf user\nhttps://guillaumegenthial.github.io/image-to-latex.html\n(check both part 1 and 2 of the web article)\n\n![](https://i.ibb.co/stbvcQY/Selection-040.png)",
    "1264858": "I have tried this, it is quite difficult on the full dataset but seems like a really interesting way to go. I’m hoping I will have some time to work on it some more before this competition is over.",
    "1265226": "i wonder if the followings make sense:\n- train a normal seq predictor\n- now modify the ground truth, e.g. reverse the sequence. train a model to predict the reverse sequence.\n\n\nif both the model gives the same results, then you know it is correct.\n\nthere are several ways that you can modify the ground truth. based on the prediction of these several modification, can we get a code that can error correct itself?\n\nthe idea is error correction of transmission codes in noisy channels, which is common in communication theory.\n\ne.g.\nhttps://sandipanweb.wordpress.com/2017/05/06/some-nlp-spelling-correction-with-noisy-channel-model/\n\neg  druchok_design_correction_control.pdf\n\n2.6 Error correction model\nError correction in SMILES strings can be framed as a standard sequence-to-sequence learning problem,\nwhich is traditionally solved with attention-based encoder-decoder models\n\nimagine if you use GAN or autoencoder to synthesis chemical formula, i wonder how do one ensure that the generated sample is of a valid molecule?",
    "1265241": "is an ensemble of several seq2seq models = another seq2se model",
    "1265308": "There was a mention of [MolGAN](https://arxiv.org/abs/1805.11973) in another discussion thread (couldn't find it anymore).\n\ntldr: MolGAN was able to generate valid molecule with high probability. Maybe one can leverage the idea in MolGAN to regularize GAN or autoencoder to generate valid molecule.",
    "1265489": "an interesting way to decode:\nhttps://iconictranslation.com/2020/05/issue-82-constrained-decoding-using-levenshtein-transformer/\n\nlevenshtein-transformer\n![](https://iconictranslation.com/wp-content/uploads/2020/05/NMT-82-LVT-fig-1.jpg)",
    "1266690": "test-time augmentation\n![](https://i.ibb.co/tYnNsNV/Selection-050.png)",
    "1269818": "![](https://miro.medium.com/max/875/0*l0J-2hNRpTBRRYQl)\n\ni almost forget there is this image bert that you can do multi-task training\n\nwe now have a unique problem:\n- we have image,text pair\n- we have additional images without text (test data)\n- we have additional text without images (external csv)\n\nhow to use all data for pretraining?",
    "1269822": "best of both worlds?\nDeepMind, Microsoft, Allen AI & UW Researchers Convert Pretrained Transformers into RNNs, Lowering Memory Cost While Retaining High Accuracy\n\nhttps://arxiv.org/pdf/2103.13076.pdf\nwould use this if it is a code challenge",
    "1269823": "MOCO-v3\n\n\"This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know\nbaseline given the recent progress in computer vision: self-supervised learning for Visual Transformers (ViT).\"\n\n\"We also verify our models in GPUs using PyTorch. It takes 24 hours for ViT-B in 128 GPUs (vs. 2.1 hours in 256\nTPUs).\"\n\nmore here:\nhttps://github.com/dk-liang/Awesome-Visual-Transformer\n\n---",
    "1269826": "CPTR: FULL TRANSFORMER NETWORK FOR IMAGE CAPTIONING\n![](https://pbs.twimg.com/media/EstJICeXEAMnt6y.jpg)",
    "1269829": "![](https://github.com/Sara-Ahmed/SiT/raw/main/SiT.png)\n\nSiT: Self-supervised image Transformer",
    "1269842": "something like this\n\n![](https://i.ibb.co/KFqdVkp/Selection-098.png)",
    "1269844": "![](https://i.ibb.co/sF077j1/Selection-101.png)",
    "1270010": "![](https://i.ibb.co/KXpfd5z/Selection-104.png)",
    "1270016": "is there a way to augment by combing 2 small molecules to a larger one?",
    "1270090": "I used an autoencoder to synthesise novel SMILES during my Masters dissertation. By sampling the latent space around known molecules it was possible to generate novel molecules. The closer the sampling to the known molecule, the more valid SMILES were produced, but at a lower uniqueness rate. \n\nNote that it is possible to write a chemical formula (e.g. SMILES or InChI) which is completely valid, but the molecule is unlikely to exist or be stable in the physical world.",
    "1270283": "![](https://github.com/clovaai/SATRN/raw/master/figures/architecture.png)\nOn Recognizing Texts of Arbitrary Shapes with 2D Self-Attention\n\n\"Adaptive 2D positional encoding (A2DPE) This new positional encoding is necessary for dynamically adapting to\nthe inherent aspect ratios incurred by overall text alignment (horizontal, diagonal, or vertical). As alternative options, we consider not doing any positional encoding at all (“None”) (Zhang et al. 2019; Wang et al. 2017), using\n1D positional encoding over flattened feature map (“1DFlatten”), using concatenation of height and width positional encodings (“2D-Concat”) (Parmar et al. 2018), and the A2DPE that we propose. See Table 3a for the results.\nWe observe that A2DPE provides the best accuracy among four options considered\"\n\n\nthis reminds me of the \"Spatial Transformer Networks\" which is used to align input using affine transformation.\n\nhere we can learn a 2d positional encoding to make input image affine invariant",
    "1270345": "i realise that you actually don't have to input the image. you can just input the non-empy block.\nhere you will have a variable length encoder.",
    "1270358": "Thanks for share! :)",
    "1270366": "![](https://i.ibb.co/CQ980Hv/Selection-106.png)\n\nkinda of error correction?\n(the input is noisy of course)",
    "1270714": "good shares",
    "1271925": "![](https://i.ibb.co/55DJqBc/Selection-029.png)\ntransformer graph decoder\n\nhttps://arxiv.org/pdf/2006.05213.pdf",
    "1273003": "graph-based encoder\n\n![](https://i.ibb.co/m5c9TBJ/Selection-032.png)",
    "1273456": "https://molvec.ncats.io/\nhttps://github.com/ncats/molvec-web\n\n\nthis should be combined with a transformer\n\n![](https://i.ibb.co/x6j7Hjp/Selection-031.png)",
    "1275982": "https://davide-belli.github.io/generative-graph-transformer.html\n![](https://davide-belli.github.io/images/blog-ggt/architecture.png)",
    "1292458": "going multi-task\n![](https://i.ibb.co/qyBj8cL/Selection-092.png)",
    "1292953": "multi-task learning is probably \n\"The Feynman Technique\"\nhttps://www.youtube.com/watch?v=6QUjy0B3lEw\n\n为什么费曼技巧被称为终极学习法 (Why feynman technique works)\nhttps://www.youtube.com/watch?v=7iNJyEbYDdc\n\n\"let the machine teach the human and not human teaches machine\"\nknowledge distillation should benefit the teacher and not the student ...\n\n---\n\nin multi-task we care about input encoding. we care less about the task.\n\nwe are forced to seek simple features (information compression) that can solve many tasks ...",
    "1295834": "How do think MLP-mixer for this application? Can it replace Transfomer? Thanks.\nhttps://arxiv.org/pdf/2105.01601.pdf",
    "1303204": "https://blogs.biomedcentral.com/on-physicalsciences/2020/03/17/how-well-can-molecules-be-generated-by-ai/\n\n\n![](https://blogs.biomedcentral.com/on-physicalsciences/wp-content/uploads/sites/14/2020/02/JCheminf_figure_1_rnn_sampling.gif)"
  },
  "source": "meta"
}