{
  "id": 243478,
  "title": "Attention is What You Get",
  "url": "/competitions/bms-molecular-translation/discussion/243478",
  "author_name": "",
  "post_date": "2021-06-02T17:49:25.311236Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Is attention really all you need?</strong> Enable/Disable CNN text feature extraction before the decoder self-attention; Increase model parameters without harming inference speed using decoder heads in series; and Experiment with my trainable &amp; parallelizable alternative to beam search.</p>\n<p>It looks like Kagglers have coalesced around \"Attention is What You Need\" models, so I have created a version that include these (novel?) features! </p>\n<p><em>Note: please use this with Google Colab. Model's \"session.run()\" calls are not working on Kaggle TPU for some unknown reason. Everything works on Colab.</em></p>\n<p><strong>My full end-to-end code: <a href=\"https://www.kaggle.com/mvenou/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/mvenou/bms-molecular-translation</a>.</strong></p>\n<hr>\n<p>Author: </p>\n<p>Mo Venouziou</p>\n<ul>\n<li><em>Email: mvenouziou@gmail.com</em></li>\n<li>*LinkedIn: <a href=\"http://www.linkedin.com/in/movenouziou/*\" target=\"_blank\">www.linkedin.com/in/movenouziou/*</a></li>\n</ul>\n<p>Updates:</p>\n<ul>\n<li><em>Original Posting: June 2, 2021</em></li>\n<li><em>06/21/21: added TPU support on Google Colab. (\"session.run()\" calls not yet working on Kaggle's TPU.)</em></li>\n<li><em>06/17/21: achieved proper training &amp; inference speed with model size matching Attention is All You Need paper on Google Colab</em></li>\n</ul>\n<hr>\n<h2>MODEL STRUCTURE:</h2>\n<p><strong>Image CNN + Attention Features encoder --&gt; text Attention + (optional )CNN feature layer decoder.</strong></p>\n<p>This is a hybrid approach with:</p>\n<ul>\n<li><p>Image Encoder from <a href=\"https://proceedings.mlr.press/v37/xuc15.pdf\" target=\"_blank\"><em>Show, Attend and Tell: Neural Image Caption Generation with Visual Attention</em></a>.  Generate image feature vectors using intermediate layer outputs from a pretrained CNN. (Here I use the more modern EfficientNet model (recommended by <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook\" target=\"_blank\"><em>Darien Schettler</em></a>) with fixed weights and a trainable Dense layer for customization.)</p></li>\n<li><p>T2T encoder-decoder model from <a href=\"https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf\" target=\"_blank\"><em>All You Need is Attention</em></a> (Self-attention feature extraction for both encoder and decoder, joint encoder-decoder attention feature interactions, and a dense prediction output block. Includes parameters to control number of encoder / decoder blocks.</p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Decoder Output Blocks placed in Series (not stacked). Increase the number of trainable parameters without adding inference computational complexity, while also allowing decoders to specialize on different regions of the output. (Note: Training is a bit trickier. My experiments show it is best to first train with the decoders using shared weights, then allowing them to vary later on in training.)</p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Is attention really all you need? Add a convolutional layer to enhance text features before decoder self-attention to experiment with performance differences with and without extra convolutional layer(s). Use of CNN's in NLP comes from <a href=\"http://proceedings.mlr.press/v70/gehring17a.html.\" target=\"_blank\"><em>Convolutional Sequence to Sequence Learning</em></a></p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Beam-Search Alternative, an extra decoding layer applied after the full logits prediction has been made. This takes the form of a bidirectional RNN applied to the full logits sequence. Because a full (initial) prediction has already been made, computations can be parallelized using statefull RNNs. (See more details below.)</p></li>\n</ul>\n<p><em>Optional features can be enabled/disabled using parameters in my model definitions.</em></p>\n<hr>\n<h2>NEXT STEPS:</h2>\n<ul>\n<li><p>(Low priority, specific to Kaggle's TPU implementation.) Fix \"session.run()\" TPU calls on Kaggle. (It works correctly on Colab.) This severely impacts inference speed on Kaggle.</p></li>\n<li><p>experiment with <strong>\"Tokens-to-Token ViT\"</strong> in place of the image CNN. (Technique from <a href=\"https://arxiv.org/pdf/2101.11986.pdf\" target=\"_blank\"><em>Training Vision Transformers from Scratch on ImageNet</em></a></p></li>\n<li><p>Train my <strong>Beam-search Alternative</strong>. </p>\n<ul>\n<li><p>Beam search is a technique to modify model predictions to reflect the (local) maximum likelihood estimate. However, it is <em>very</em> local in that computation expense increases quickly with the number of character steps taken into account. This is also a hard-coded algorithm, which is somewhat contrary to the philosophy of deep learning.</p></li>\n<li><p>A <em>Beam-search Alternative</em> would be an extra decoding layer applied <em>after</em> the full logits prediction has been made. This might be in the form of a stateful, bidirectional RNN that is computationally parallizable because it is applied to the full logits sequence.</p></li>\n<li><p>Need to revamp code to accept main model changes made for TPU support.</p></li></ul></li>\n<li><p>Treat the number of convolutional layers (decoder feature extraction) and number of decoders places in series (decoder prediction output) as <strong>new hyperparameters</strong> to tune.</p></li>\n<li><p><em>6/21/21: TPU Support added on Colab</em> </p></li>\n<li><p>*6/17/21: Increased model size and efficiency. * </p></li>\n</ul>\n<hr>\n<h3>CITATIONS</h3>\n<ul>\n<li><p>\"Attention is All You Need.\" </p>\n<ul>\n<li>Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin. NIPS (2017). <em><a href=\"https://research.google/pubs/pub46201/\" target=\"_blank\">https://research.google/pubs/pub46201/</a></em></li></ul></li>\n<li><p>\"Convolutional Sequence to Sequence Learning.\"</p>\n<ul>\n<li>Gehring, J., Auli, M., Grangier, D., Yarats, D. &amp; Dauphin, Y.N.. (2017). Convolutional Sequence to Sequence Learning. Proceedings of the 34th International Conference on Machine Learning, in Proceedings of Machine Learning Research 70:1243-1252, *<a href=\"http://proceedings.mlr.press/v70/gehring17a.html.*\" target=\"_blank\">http://proceedings.mlr.press/v70/gehring17a.html.*</a></li></ul></li>\n<li><p>\"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\"</p>\n<ul>\n<li>Mingxing Tan, Quoc V. Le (2019). Convolutional Sequence to Sequence Learning. International Conference on Machine Learning. *<a href=\"http://arxiv.org/abs/1905.11946.*\" target=\"_blank\">http://arxiv.org/abs/1905.11946.*</a></li></ul></li>\n<li><p>\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.\"</p>\n<ul>\n<li>Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R. &amp; Bengio, Y.. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the 32nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 37:2048-2057. *<a href=\"http://proceedings.mlr.press/v37/xuc15.html.*\" target=\"_blank\">http://proceedings.mlr.press/v37/xuc15.html.*</a> </li></ul></li>\n<li><p>\"Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet\"</p>\n<ul>\n<li>Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, Shuicheng Yan. Preprint (2021). *<a href=\"https://arxiv.org/abs/2101.11986*\" target=\"_blank\">https://arxiv.org/abs/2101.11986*</a>.</li></ul></li>\n<li><p>Tensorflow documentation tutorial \"Transformer model for language understanding.\" I found this after fully completing the model and found the attention mask was incorrect. My use of \"tf.linalg.band_part\" (only) is due to this tutorial. <em><a href=\"http://www.tensorflow.org/text/tutorials/transformer#masking\" target=\"_blank\">www.tensorflow.org/text/tutorials/transformer#masking</a></em></p></li>\n<li><p>Special thanks to <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook.\" target=\"_blank\">Darien Schettler</a> for leading readers to the \"Show\" and \"Attention\" papers cited above, using <em>session.run()</em> to improve TPU inference speed in distributed settings and providing detailed info on creating TF Records. This work is otherwise derived independently from his.</p></li>\n<li><p>It is possible my idea of a Beam Search Alternative is based on a lecture video from DeepLearning.ai's <a href=\"https://www.coursera.org/specializations/deep-learning\" target=\"_blank\">Deep Learning Specialization</a>  on Coursera.</p></li>\n<li><p><strong>Dataset / Kaggle Competition:</strong> \"Bristol-Myers Squibb – Molecular Translation\" competition on Kaggle (2021). <em><a href=\"https://www.kaggle.com/c/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation</a></em></p></li>\n</ul>\n<hr>",
  "messages": [
    {
      "id": "1333427",
      "postDate": "06/02/2021 17:49:25",
      "content": "<p><strong>Is attention really all you need?</strong> Enable/Disable CNN text feature extraction before the decoder self-attention; Increase model parameters without harming inference speed using decoder heads in series; and Experiment with my trainable &amp; parallelizable alternative to beam search.</p>\n<p>It looks like Kagglers have coalesced around \"Attention is What You Need\" models, so I have created a version that include these (novel?) features! </p>\n<p><em>Note: please use this with Google Colab. Model's \"session.run()\" calls are not working on Kaggle TPU for some unknown reason. Everything works on Colab.</em></p>\n<p><strong>My full end-to-end code: <a href=\"https://www.kaggle.com/mvenou/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/mvenou/bms-molecular-translation</a>.</strong></p>\n<hr>\n<p>Author: </p>\n<p>Mo Venouziou</p>\n<ul>\n<li><em>Email: mvenouziou@gmail.com</em></li>\n<li>*LinkedIn: <a href=\"http://www.linkedin.com/in/movenouziou/*\" target=\"_blank\">www.linkedin.com/in/movenouziou/*</a></li>\n</ul>\n<p>Updates:</p>\n<ul>\n<li><em>Original Posting: June 2, 2021</em></li>\n<li><em>06/21/21: added TPU support on Google Colab. (\"session.run()\" calls not yet working on Kaggle's TPU.)</em></li>\n<li><em>06/17/21: achieved proper training &amp; inference speed with model size matching Attention is All You Need paper on Google Colab</em></li>\n</ul>\n<hr>\n<h2>MODEL STRUCTURE:</h2>\n<p><strong>Image CNN + Attention Features encoder --&gt; text Attention + (optional )CNN feature layer decoder.</strong></p>\n<p>This is a hybrid approach with:</p>\n<ul>\n<li><p>Image Encoder from <a href=\"https://proceedings.mlr.press/v37/xuc15.pdf\" target=\"_blank\"><em>Show, Attend and Tell: Neural Image Caption Generation with Visual Attention</em></a>.  Generate image feature vectors using intermediate layer outputs from a pretrained CNN. (Here I use the more modern EfficientNet model (recommended by <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook\" target=\"_blank\"><em>Darien Schettler</em></a>) with fixed weights and a trainable Dense layer for customization.)</p></li>\n<li><p>T2T encoder-decoder model from <a href=\"https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf\" target=\"_blank\"><em>All You Need is Attention</em></a> (Self-attention feature extraction for both encoder and decoder, joint encoder-decoder attention feature interactions, and a dense prediction output block. Includes parameters to control number of encoder / decoder blocks.</p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Decoder Output Blocks placed in Series (not stacked). Increase the number of trainable parameters without adding inference computational complexity, while also allowing decoders to specialize on different regions of the output. (Note: Training is a bit trickier. My experiments show it is best to first train with the decoders using shared weights, then allowing them to vary later on in training.)</p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Is attention really all you need? Add a convolutional layer to enhance text features before decoder self-attention to experiment with performance differences with and without extra convolutional layer(s). Use of CNN's in NLP comes from <a href=\"http://proceedings.mlr.press/v70/gehring17a.html.\" target=\"_blank\"><em>Convolutional Sequence to Sequence Learning</em></a></p></li>\n<li><p><strong><em>PLUS</em></strong> <em>(optional):</em> Beam-Search Alternative, an extra decoding layer applied after the full logits prediction has been made. This takes the form of a bidirectional RNN applied to the full logits sequence. Because a full (initial) prediction has already been made, computations can be parallelized using statefull RNNs. (See more details below.)</p></li>\n</ul>\n<p><em>Optional features can be enabled/disabled using parameters in my model definitions.</em></p>\n<hr>\n<h2>NEXT STEPS:</h2>\n<ul>\n<li><p>(Low priority, specific to Kaggle's TPU implementation.) Fix \"session.run()\" TPU calls on Kaggle. (It works correctly on Colab.) This severely impacts inference speed on Kaggle.</p></li>\n<li><p>experiment with <strong>\"Tokens-to-Token ViT\"</strong> in place of the image CNN. (Technique from <a href=\"https://arxiv.org/pdf/2101.11986.pdf\" target=\"_blank\"><em>Training Vision Transformers from Scratch on ImageNet</em></a></p></li>\n<li><p>Train my <strong>Beam-search Alternative</strong>. </p>\n<ul>\n<li><p>Beam search is a technique to modify model predictions to reflect the (local) maximum likelihood estimate. However, it is <em>very</em> local in that computation expense increases quickly with the number of character steps taken into account. This is also a hard-coded algorithm, which is somewhat contrary to the philosophy of deep learning.</p></li>\n<li><p>A <em>Beam-search Alternative</em> would be an extra decoding layer applied <em>after</em> the full logits prediction has been made. This might be in the form of a stateful, bidirectional RNN that is computationally parallizable because it is applied to the full logits sequence.</p></li>\n<li><p>Need to revamp code to accept main model changes made for TPU support.</p></li></ul></li>\n<li><p>Treat the number of convolutional layers (decoder feature extraction) and number of decoders places in series (decoder prediction output) as <strong>new hyperparameters</strong> to tune.</p></li>\n<li><p><em>6/21/21: TPU Support added on Colab</em> </p></li>\n<li><p>*6/17/21: Increased model size and efficiency. * </p></li>\n</ul>\n<hr>\n<h3>CITATIONS</h3>\n<ul>\n<li><p>\"Attention is All You Need.\" </p>\n<ul>\n<li>Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin. NIPS (2017). <em><a href=\"https://research.google/pubs/pub46201/\" target=\"_blank\">https://research.google/pubs/pub46201/</a></em></li></ul></li>\n<li><p>\"Convolutional Sequence to Sequence Learning.\"</p>\n<ul>\n<li>Gehring, J., Auli, M., Grangier, D., Yarats, D. &amp; Dauphin, Y.N.. (2017). Convolutional Sequence to Sequence Learning. Proceedings of the 34th International Conference on Machine Learning, in Proceedings of Machine Learning Research 70:1243-1252, *<a href=\"http://proceedings.mlr.press/v70/gehring17a.html.*\" target=\"_blank\">http://proceedings.mlr.press/v70/gehring17a.html.*</a></li></ul></li>\n<li><p>\"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\"</p>\n<ul>\n<li>Mingxing Tan, Quoc V. Le (2019). Convolutional Sequence to Sequence Learning. International Conference on Machine Learning. *<a href=\"http://arxiv.org/abs/1905.11946.*\" target=\"_blank\">http://arxiv.org/abs/1905.11946.*</a></li></ul></li>\n<li><p>\"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.\"</p>\n<ul>\n<li>Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R. &amp; Bengio, Y.. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the 32nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 37:2048-2057. *<a href=\"http://proceedings.mlr.press/v37/xuc15.html.*\" target=\"_blank\">http://proceedings.mlr.press/v37/xuc15.html.*</a> </li></ul></li>\n<li><p>\"Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet\"</p>\n<ul>\n<li>Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, Shuicheng Yan. Preprint (2021). *<a href=\"https://arxiv.org/abs/2101.11986*\" target=\"_blank\">https://arxiv.org/abs/2101.11986*</a>.</li></ul></li>\n<li><p>Tensorflow documentation tutorial \"Transformer model for language understanding.\" I found this after fully completing the model and found the attention mask was incorrect. My use of \"tf.linalg.band_part\" (only) is due to this tutorial. <em><a href=\"http://www.tensorflow.org/text/tutorials/transformer#masking\" target=\"_blank\">www.tensorflow.org/text/tutorials/transformer#masking</a></em></p></li>\n<li><p>Special thanks to <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook.\" target=\"_blank\">Darien Schettler</a> for leading readers to the \"Show\" and \"Attention\" papers cited above, using <em>session.run()</em> to improve TPU inference speed in distributed settings and providing detailed info on creating TF Records. This work is otherwise derived independently from his.</p></li>\n<li><p>It is possible my idea of a Beam Search Alternative is based on a lecture video from DeepLearning.ai's <a href=\"https://www.coursera.org/specializations/deep-learning\" target=\"_blank\">Deep Learning Specialization</a>  on Coursera.</p></li>\n<li><p><strong>Dataset / Kaggle Competition:</strong> \"Bristol-Myers Squibb – Molecular Translation\" competition on Kaggle (2021). <em><a href=\"https://www.kaggle.com/c/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation</a></em></p></li>\n</ul>\n<hr>",
      "rawMarkdown": "**Is attention really all you need?** Enable/Disable CNN text feature extraction before the decoder self-attention; Increase model parameters without harming inference speed using decoder heads in series; and Experiment with my trainable & parallelizable alternative to beam search.\n\nIt looks like Kagglers have coalesced around \"Attention is What You Need\" models, so I have created a version that include these (novel?) features! \n\n*Note: please use this with Google Colab. Model's \"session.run()\" calls are not working on Kaggle TPU for some unknown reason. Everything works on Colab.*\n\n**My full end-to-end code: [https://www.kaggle.com/mvenou/bms-molecular-translation](https://www.kaggle.com/mvenou/bms-molecular-translation).**\n\n-----\n\nAuthor: \n\nMo Venouziou\n\n- *Email: mvenouziou@gmail.com*\n- *LinkedIn: www.linkedin.com/in/movenouziou/*\n\nUpdates:\n\n - *Original Posting: June 2, 2021*\n - *06/21/21: added TPU support on Google Colab. (\"session.run()\" calls not yet working on Kaggle's TPU.)*\n - *06/17/21: achieved proper training & inference speed with model size matching Attention is All You Need paper on Google Colab*\n\n\n----\n\n## MODEL STRUCTURE: \n\n**Image CNN + Attention Features encoder --> text Attention + (optional )CNN feature layer decoder.**\n\nThis is a hybrid approach with:\n \n - Image Encoder from [*Show, Attend and Tell: Neural Image Caption Generation with Visual Attention*](https://proceedings.mlr.press/v37/xuc15.pdf).  Generate image feature vectors using intermediate layer outputs from a pretrained CNN. (Here I use the more modern EfficientNet model (recommended by [*Darien Schettler*](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook)) with fixed weights and a trainable Dense layer for customization.)\n \n - T2T encoder-decoder model from [*All You Need is Attention*](https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf) (Self-attention feature extraction for both encoder and decoder, joint encoder-decoder attention feature interactions, and a dense prediction output block. Includes parameters to control number of encoder / decoder blocks.\n\n - ***PLUS*** *(optional):* Decoder Output Blocks placed in Series (not stacked). Increase the number of trainable parameters without adding inference computational complexity, while also allowing decoders to specialize on different regions of the output. (Note: Training is a bit trickier. My experiments show it is best to first train with the decoders using shared weights, then allowing them to vary later on in training.)\n \n - ***PLUS*** *(optional):* Is attention really all you need? Add a convolutional layer to enhance text features before decoder self-attention to experiment with performance differences with and without extra convolutional layer(s). Use of CNN's in NLP comes from [*Convolutional Sequence to Sequence Learning*](http://proceedings.mlr.press/v70/gehring17a.html.)\n\n - ***PLUS*** *(optional):* Beam-Search Alternative, an extra decoding layer applied after the full logits prediction has been made. This takes the form of a bidirectional RNN applied to the full logits sequence. Because a full (initial) prediction has already been made, computations can be parallelized using statefull RNNs. (See more details below.)\n\n*Optional features can be enabled/disabled using parameters in my model definitions.*\n\n----\n\n## NEXT STEPS:\n\n - (Low priority, specific to Kaggle's TPU implementation.) Fix \"session.run()\" TPU calls on Kaggle. (It works correctly on Colab.) This severely impacts inference speed on Kaggle.\n\n - experiment with **\"Tokens-to-Token ViT\"** in place of the image CNN. (Technique from [*Training Vision Transformers from Scratch on ImageNet*](https://arxiv.org/pdf/2101.11986.pdf)\n  \n - Train my **Beam-search Alternative**. \n\n    - Beam search is a technique to modify model predictions to reflect the (local) maximum likelihood estimate. However, it is *very* local in that computation expense increases quickly with the number of character steps taken into account. This is also a hard-coded algorithm, which is somewhat contrary to the philosophy of deep learning.\n\n    - A *Beam-search Alternative* would be an extra decoding layer applied *after* the full logits prediction has been made. This might be in the form of a stateful, bidirectional RNN that is computationally parallizable because it is applied to the full logits sequence.\n\n    - Need to revamp code to accept main model changes made for TPU support.\n\n - Treat the number of convolutional layers (decoder feature extraction) and number of decoders places in series (decoder prediction output) as **new hyperparameters** to tune.\n\n - *6/21/21: TPU Support added on Colab* ~~Implement TPU training. (Currently runs on GPU. Note that a CPU alone is not enough to achieve acceptable inference speed.)~~\n\n - *6/17/21: Increased model size and efficiency. * ~~ Implement full size model (matching AISYN) with efficient training and inference speeds for the large dataset. (TPU required. GPU doesn't have enough memory to train such a large model)~~\n\n----\n\n\n### CITATIONS\n\n- \"Attention is All You Need.\" \n - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin. NIPS (2017). *https://research.google/pubs/pub46201/*\n\n- \"Convolutional Sequence to Sequence Learning.\"\n \n  - Gehring, J., Auli, M., Grangier, D., Yarats, D. & Dauphin, Y.N.. (2017). Convolutional Sequence to Sequence Learning. Proceedings of the 34th International Conference on Machine Learning, in Proceedings of Machine Learning Research 70:1243-1252, *http://proceedings.mlr.press/v70/gehring17a.html.*\n\n\n- \"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\"\n \n  - Mingxing Tan, Quoc V. Le (2019). Convolutional Sequence to Sequence Learning. International Conference on Machine Learning. *http://arxiv.org/abs/1905.11946.*\n\n\n-  \"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.\"\n  -  Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R. & Bengio, Y.. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the 32nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 37:2048-2057. *http://proceedings.mlr.press/v37/xuc15.html.* \n            \n\n- \"Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet\"\n\n  - Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, Shuicheng Yan. Preprint (2021). *https://arxiv.org/abs/2101.11986*.\n\n- Tensorflow documentation tutorial \"Transformer model for language understanding.\" I found this after fully completing the model and found the attention mask was incorrect. My use of \"tf.linalg.band_part\" (only) is due to this tutorial. *www.tensorflow.org/text/tutorials/transformer#masking*\n\n- Special thanks to [Darien Schettler](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook.) for leading readers to the \"Show\" and \"Attention\" papers cited above, using *session.run()* to improve TPU inference speed in distributed settings and providing detailed info on creating TF Records. This work is otherwise derived independently from his.\n\n- It is possible my idea of a Beam Search Alternative is based on a lecture video from DeepLearning.ai's [Deep Learning Specialization](https://www.coursera.org/specializations/deep-learning)  on Coursera.\n\n- **Dataset / Kaggle Competition:** \"Bristol-Myers Squibb – Molecular Translation\" competition on Kaggle (2021). *https://www.kaggle.com/c/bms-molecular-translation*\n\n----",
      "votes": null
    },
    {
      "id": "1334323",
      "postDate": "06/03/2021 12:21:58",
      "content": "<p>Looks great! Excited to read through and learn.</p>\n<p>Just an FYI though, your code link is going to the general competition page not a notebook… I assume this is a typo/error.</p>",
      "rawMarkdown": "Looks great! Excited to read through and learn.\n\nJust an FYI though, your code link is going to the general competition page not a notebook... I assume this is a typo/error.",
      "votes": null
    },
    {
      "id": "1334662",
      "postDate": "06/03/2021 17:03:42",
      "content": "<p>Thanks for catching that-- link now updated!</p>",
      "rawMarkdown": "Thanks for catching that-- link now updated!",
      "votes": null
    },
    {
      "id": "1354811",
      "postDate": "06/17/2021 22:21:52",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thank you for posting many of your notebooks throughout this competition. I was having a lot of difficulty getting the TPU to work correctly / efficiently and the notebooks were of great help!</p>",
      "rawMarkdown": "dschettler8845, thank you for posting many of your notebooks throughout this competition. I was having a lot of difficulty getting the TPU to work correctly / efficiently and the notebooks were of great help!",
      "votes": null
    },
    {
      "id": "1354824",
      "postDate": "06/17/2021 22:59:24",
      "content": "<p>Happy to help !</p>",
      "rawMarkdown": "Happy to help !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1334323,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/03/2021 12:21:58",
      "content": "<p>Looks great! Excited to read through and learn.</p>\n<p>Just an FYI though, your code link is going to the general competition page not a notebook… I assume this is a typo/error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334662,
          "author_name": "mvenou",
          "author_url": "",
          "post_date": "06/03/2021 17:03:42",
          "content": "<p>Thanks for catching that-- link now updated!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1354811,
          "author_name": "mvenou",
          "author_url": "",
          "post_date": "06/17/2021 22:21:52",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thank you for posting many of your notebooks throughout this competition. I was having a lot of difficulty getting the TPU to work correctly / efficiently and the notebooks were of great help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1354824,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "06/17/2021 22:59:24",
          "content": "<p>Happy to help !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1333427": "**Is attention really all you need?** Enable/Disable CNN text feature extraction before the decoder self-attention; Increase model parameters without harming inference speed using decoder heads in series; and Experiment with my trainable & parallelizable alternative to beam search.\n\nIt looks like Kagglers have coalesced around \"Attention is What You Need\" models, so I have created a version that include these (novel?) features! \n\n*Note: please use this with Google Colab. Model's \"session.run()\" calls are not working on Kaggle TPU for some unknown reason. Everything works on Colab.*\n\n**My full end-to-end code: [https://www.kaggle.com/mvenou/bms-molecular-translation](https://www.kaggle.com/mvenou/bms-molecular-translation).**\n\n-----\n\nAuthor: \n\nMo Venouziou\n\n- *Email: mvenouziou@gmail.com*\n- *LinkedIn: www.linkedin.com/in/movenouziou/*\n\nUpdates:\n\n - *Original Posting: June 2, 2021*\n - *06/21/21: added TPU support on Google Colab. (\"session.run()\" calls not yet working on Kaggle's TPU.)*\n - *06/17/21: achieved proper training & inference speed with model size matching Attention is All You Need paper on Google Colab*\n\n\n----\n\n## MODEL STRUCTURE: \n\n**Image CNN + Attention Features encoder --> text Attention + (optional )CNN feature layer decoder.**\n\nThis is a hybrid approach with:\n \n - Image Encoder from [*Show, Attend and Tell: Neural Image Caption Generation with Visual Attention*](https://proceedings.mlr.press/v37/xuc15.pdf).  Generate image feature vectors using intermediate layer outputs from a pretrained CNN. (Here I use the more modern EfficientNet model (recommended by [*Darien Schettler*](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook)) with fixed weights and a trainable Dense layer for customization.)\n \n - T2T encoder-decoder model from [*All You Need is Attention*](https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf) (Self-attention feature extraction for both encoder and decoder, joint encoder-decoder attention feature interactions, and a dense prediction output block. Includes parameters to control number of encoder / decoder blocks.\n\n - ***PLUS*** *(optional):* Decoder Output Blocks placed in Series (not stacked). Increase the number of trainable parameters without adding inference computational complexity, while also allowing decoders to specialize on different regions of the output. (Note: Training is a bit trickier. My experiments show it is best to first train with the decoders using shared weights, then allowing them to vary later on in training.)\n \n - ***PLUS*** *(optional):* Is attention really all you need? Add a convolutional layer to enhance text features before decoder self-attention to experiment with performance differences with and without extra convolutional layer(s). Use of CNN's in NLP comes from [*Convolutional Sequence to Sequence Learning*](http://proceedings.mlr.press/v70/gehring17a.html.)\n\n - ***PLUS*** *(optional):* Beam-Search Alternative, an extra decoding layer applied after the full logits prediction has been made. This takes the form of a bidirectional RNN applied to the full logits sequence. Because a full (initial) prediction has already been made, computations can be parallelized using statefull RNNs. (See more details below.)\n\n*Optional features can be enabled/disabled using parameters in my model definitions.*\n\n----\n\n## NEXT STEPS:\n\n - (Low priority, specific to Kaggle's TPU implementation.) Fix \"session.run()\" TPU calls on Kaggle. (It works correctly on Colab.) This severely impacts inference speed on Kaggle.\n\n - experiment with **\"Tokens-to-Token ViT\"** in place of the image CNN. (Technique from [*Training Vision Transformers from Scratch on ImageNet*](https://arxiv.org/pdf/2101.11986.pdf)\n  \n - Train my **Beam-search Alternative**. \n\n    - Beam search is a technique to modify model predictions to reflect the (local) maximum likelihood estimate. However, it is *very* local in that computation expense increases quickly with the number of character steps taken into account. This is also a hard-coded algorithm, which is somewhat contrary to the philosophy of deep learning.\n\n    - A *Beam-search Alternative* would be an extra decoding layer applied *after* the full logits prediction has been made. This might be in the form of a stateful, bidirectional RNN that is computationally parallizable because it is applied to the full logits sequence.\n\n    - Need to revamp code to accept main model changes made for TPU support.\n\n - Treat the number of convolutional layers (decoder feature extraction) and number of decoders places in series (decoder prediction output) as **new hyperparameters** to tune.\n\n - *6/21/21: TPU Support added on Colab* ~~Implement TPU training. (Currently runs on GPU. Note that a CPU alone is not enough to achieve acceptable inference speed.)~~\n\n - *6/17/21: Increased model size and efficiency. * ~~ Implement full size model (matching AISYN) with efficient training and inference speeds for the large dataset. (TPU required. GPU doesn't have enough memory to train such a large model)~~\n\n----\n\n\n### CITATIONS\n\n- \"Attention is All You Need.\" \n - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin. NIPS (2017). *https://research.google/pubs/pub46201/*\n\n- \"Convolutional Sequence to Sequence Learning.\"\n \n  - Gehring, J., Auli, M., Grangier, D., Yarats, D. & Dauphin, Y.N.. (2017). Convolutional Sequence to Sequence Learning. Proceedings of the 34th International Conference on Machine Learning, in Proceedings of Machine Learning Research 70:1243-1252, *http://proceedings.mlr.press/v70/gehring17a.html.*\n\n\n- \"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\"\n \n  - Mingxing Tan, Quoc V. Le (2019). Convolutional Sequence to Sequence Learning. International Conference on Machine Learning. *http://arxiv.org/abs/1905.11946.*\n\n\n-  \"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.\"\n  -  Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R. & Bengio, Y.. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the 32nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 37:2048-2057. *http://proceedings.mlr.press/v37/xuc15.html.* \n            \n\n- \"Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet\"\n\n  - Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, Shuicheng Yan. Preprint (2021). *https://arxiv.org/abs/2101.11986*.\n\n- Tensorflow documentation tutorial \"Transformer model for language understanding.\" I found this after fully completing the model and found the attention mask was incorrect. My use of \"tf.linalg.band_part\" (only) is due to this tutorial. *www.tensorflow.org/text/tutorials/transformer#masking*\n\n- Special thanks to [Darien Schettler](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/notebook.) for leading readers to the \"Show\" and \"Attention\" papers cited above, using *session.run()* to improve TPU inference speed in distributed settings and providing detailed info on creating TF Records. This work is otherwise derived independently from his.\n\n- It is possible my idea of a Beam Search Alternative is based on a lecture video from DeepLearning.ai's [Deep Learning Specialization](https://www.coursera.org/specializations/deep-learning)  on Coursera.\n\n- **Dataset / Kaggle Competition:** \"Bristol-Myers Squibb – Molecular Translation\" competition on Kaggle (2021). *https://www.kaggle.com/c/bms-molecular-translation*\n\n----",
    "1334323": "Looks great! Excited to read through and learn.\n\nJust an FYI though, your code link is going to the general competition page not a notebook... I assume this is a typo/error.",
    "1334662": "Thanks for catching that-- link now updated!",
    "1354811": "dschettler8845, thank you for posting many of your notebooks throughout this competition. I was having a lot of difficulty getting the TPU to work correctly / efficiently and the notebooks were of great help!",
    "1354824": "Happy to help !"
  },
  "source": "meta"
}