{
  "id": 241716,
  "title": "Transformer Decoder + EfficientNetV2 Encoder – Full Pipeline – [3.71 CV]",
  "url": "/competitions/bms-molecular-translation/discussion/241716",
  "author_name": "",
  "post_date": "2021-05-25T19:47:00.950107900Z",
  "votes": 28,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi there, </p>\n<p>I have created a <a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\">notebook</a> that implements a pipeline consisting of a Transformer Decoder and an EfficientNetV2 Encoder (all in Tensorflow).</p>\n<p>This was a bit more difficult than I anticipated and I learned a lot (it took all 30 TPU hours this week and I still didn't get inference on the test set working). I just want to take this opportunity and share the knowledge and struggles with everyone.</p>\n<p>The full architecture was able to achieve <strong>CV=3.71 in 6 epochs</strong> <em>(although I loaded the EfficientNetV2 pre-trained model from my other pipeline as a starting point for the encoder)</em></p>\n<p>The learning rate scheduling is difficult to get right and I encountered lots and lots of trouble with spiking loss and weird stuff happening. I think batching by sequence length and using a lower max LR would help considerably.</p>\n<hr>\n<p><br></p>\n<p><strong><em>TL;DR  –––––––  TF+TPU+EfficientNetV2+Decoder  –––&gt;</em> DIFFICULT TO TRAIN … BUT 🔥🔥🧑‍🚒(me)🔥🔥</strong></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><em>A <strong>caveat</strong>: I ran out of TPU hours and was not able to get the test inference to work. Also, the loss/accuracy became unstable during my last training (LR curve should be flatter and lower probably). All this being equal, the validation inference DOES work, and it's very similar to the test inference… so I must just be missing something. But as such, I have not submitted this notebook to the public LB yet. If someone gets it to work please let me know how it does!</em></p>\n<p><strong><em>ADDENDUM</em></strong>: </p>\n<ul>\n<li>Apparently my notebook continued after my TPU ran out. I just saw that my notebook was able to START inferring on the test set (who knows if it will finish and if I can save it). It will take about 1300 seconds to infer across the whole test set and another 5-10 for post-processing. </li>\n<li>So ~30 minutes for total inference time on the test dataset using this pipeline</li>\n</ul>\n<p><strong><em>FINAL ADDENDUM</em></strong>: </p>\n<ul>\n<li>Somehow my save and submission went through. This notebook scores <strong>5.83 LB LSD</strong><br>\n<br></li>\n</ul>\n<hr>\n<p><br></p>\n<h3><b>Links</b></h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\"><strong>My End-to-End Transformer+EfficientNetV2 Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\"><strong>My Previous EfficientNetV2+LSTM End-to-End Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2\" target=\"_blank\"><strong>My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\"><strong>How To Load ImageNet Weights Into EfficientNetV2 And Infer</strong> <strong><em>(for fine-tuning)</em></strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\"><strong>The EfficientNetV2 Paper</strong></a></li>\n<li><a href=\"https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2\" target=\"_blank\"><strong>The Official Google Implementation of EfficientNetV2</strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/1502.03044\" target=\"_blank\"><strong>The Show, Attend, and Tell Paper</strong></a></li>\n<li><a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\"><strong>The Attention Is All You Need Paper</strong></a></li>\n</ul>\n<p><br></p>\n<h3><b>Model Details</b></h3>\n<ul>\n<li><strong>Encoder</strong>: <ul>\n<li><strong><em>EfficientNetV2-B2</em></strong> - Weights From <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\"><strong>Here</strong></a></li></ul></li>\n<li><strong>Decoder:</strong> <ul>\n<li><em>Transformer</em></li>\n<li><em>Number of Layers -&gt; 4</em></li>\n<li><em>Number of Heads -&gt; 8</em></li>\n<li><strong><em><code>d_model</code></em></strong> <em>-&gt; Image Embedding Depth (<strong><code>208</code></strong>)</em></li>\n<li><em>Units In Pointwise Feed-Forward NN -&gt; 512</em></li>\n<li><em>Dropout -&gt; 0.1</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Data Details</b></h3>\n<ul>\n<li><strong>Data Size:</strong><ul>\n<li><em>Val -&gt; 80,000 (random sample)</em></li>\n<li><em>Train -&gt; All-80,000</em></li>\n<li><em>Test -&gt; All</em></li></ul></li>\n<li><strong>Batch Size:</strong><ul>\n<li><em>128 Per Replica (1024 overall)</em></li></ul></li>\n<li><strong>Image Size:</strong><ul>\n<li><em>192x384x3</em></li></ul></li>\n<li><strong>Image Prep:</strong><ul>\n<li><em>TRAIN: Crop To Molecule, Invert and Resize (no padding)</em></li>\n<li><em>TEST : Crop To Molecule, Invert and Resize (no padding) … <strong>also if h&gt;w rotate 90 degrees</strong></em></li></ul></li>\n<li><strong>InChI Manipulation:</strong><ul>\n<li><em>Limit Max Length to Be Half Of Actual Max Length</em></li></ul></li>\n<li><strong>InChI Tokenization:</strong><ul>\n<li><em>Entities Are Considered Tokens (i.e. <code>Br</code> is a token, <code>191</code> is a token, etc.)</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Pipeline and General Details</b></h3>\n<ul>\n<li><strong>Execution &amp; Development Environment</strong><ul>\n<li><em>Kaggle Notebook</em></li></ul></li>\n<li><strong>Framework</strong><ul>\n<li><em>Tensorflow</em></li></ul></li>\n<li><strong>Accelerator</strong><ul>\n<li><em>TPU v3-8</em></li></ul></li>\n<li><strong>Data Format</strong><ul>\n<li><em>TFRecords</em></li></ul></li>\n<li><strong>I/O Streaming Tools</strong><ul>\n<li><em>tf.data</em></li></ul></li>\n<li><strong>Learning Rate Scheduler</strong><ul>\n<li><strong><em>Implemented from the \"Attention is All You Need Paper\" at half the magnitude… still too high. Very unstable.</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Result Details</b></h3>\n<ul>\n<li><strong>Timing</strong><ul>\n<li><em>~1.5-2.5 hours to complete training</em></li>\n<li><em>~8-12 minutes per epoch (includes validation)</em></li>\n<li><em>~40-45 minutes for the entire inference and submission process</em></li>\n<li><em>~35 minutes to infer on the test data</em></li>\n<li><em>8 minutes and 32 seconds to post-process test data (tokens to strings)</em></li>\n<li><strong><em>~2.50-4.25 hours to run the entire notebook including extras and visualizations</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<ul>\n<li><strong>Validation Levenshtein Distance (up to half max length) By Epoch (every other)</strong><ul>\n<li><em>2  -&gt; 11.55 CV LSD</em></li>\n<li><em>4  -&gt; 4.85 CV LSD</em></li>\n<li><strong><em>6  -&gt; 3.71 CV LSD</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<ul>\n<li><strong>Final Validation Levenshtein Distance</strong><ul>\n<li><strong><em>NOT COMPUTED</em></strong></li></ul></li>\n<li><strong>Public Leaderboard Score</strong><ul>\n<li><strong><em>5.83 LB LSD</em></strong></li></ul></li>\n</ul>\n<p>This was just where I had to stop because I ran out of TPU… I assume with a better schedule you could take this a lot further. Even on half-length tokens.</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p>I hope this helps! It was an effort and a half to get this to work and I'm certain that I did things that don't make sense or could be better. If you see anything like this please don't hesitate to let me know. I'm all about learning and getting better.</p>\n<p><br></p>\n<hr>\n<p>Thank you for your time!</p>",
  "messages": [
    {
      "id": "1322932",
      "postDate": "05/25/2021 19:47:00",
      "content": "<p>Hi there, </p>\n<p>I have created a <a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\">notebook</a> that implements a pipeline consisting of a Transformer Decoder and an EfficientNetV2 Encoder (all in Tensorflow).</p>\n<p>This was a bit more difficult than I anticipated and I learned a lot (it took all 30 TPU hours this week and I still didn't get inference on the test set working). I just want to take this opportunity and share the knowledge and struggles with everyone.</p>\n<p>The full architecture was able to achieve <strong>CV=3.71 in 6 epochs</strong> <em>(although I loaded the EfficientNetV2 pre-trained model from my other pipeline as a starting point for the encoder)</em></p>\n<p>The learning rate scheduling is difficult to get right and I encountered lots and lots of trouble with spiking loss and weird stuff happening. I think batching by sequence length and using a lower max LR would help considerably.</p>\n<hr>\n<p><br></p>\n<p><strong><em>TL;DR  –––––––  TF+TPU+EfficientNetV2+Decoder  –––&gt;</em> DIFFICULT TO TRAIN … BUT 🔥🔥🧑‍🚒(me)🔥🔥</strong></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><em>A <strong>caveat</strong>: I ran out of TPU hours and was not able to get the test inference to work. Also, the loss/accuracy became unstable during my last training (LR curve should be flatter and lower probably). All this being equal, the validation inference DOES work, and it's very similar to the test inference… so I must just be missing something. But as such, I have not submitted this notebook to the public LB yet. If someone gets it to work please let me know how it does!</em></p>\n<p><strong><em>ADDENDUM</em></strong>: </p>\n<ul>\n<li>Apparently my notebook continued after my TPU ran out. I just saw that my notebook was able to START inferring on the test set (who knows if it will finish and if I can save it). It will take about 1300 seconds to infer across the whole test set and another 5-10 for post-processing. </li>\n<li>So ~30 minutes for total inference time on the test dataset using this pipeline</li>\n</ul>\n<p><strong><em>FINAL ADDENDUM</em></strong>: </p>\n<ul>\n<li>Somehow my save and submission went through. This notebook scores <strong>5.83 LB LSD</strong><br>\n<br></li>\n</ul>\n<hr>\n<p><br></p>\n<h3><b>Links</b></h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\"><strong>My End-to-End Transformer+EfficientNetV2 Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\"><strong>My Previous EfficientNetV2+LSTM End-to-End Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2\" target=\"_blank\"><strong>My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\"><strong>How To Load ImageNet Weights Into EfficientNetV2 And Infer</strong> <strong><em>(for fine-tuning)</em></strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\"><strong>The EfficientNetV2 Paper</strong></a></li>\n<li><a href=\"https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2\" target=\"_blank\"><strong>The Official Google Implementation of EfficientNetV2</strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/1502.03044\" target=\"_blank\"><strong>The Show, Attend, and Tell Paper</strong></a></li>\n<li><a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\"><strong>The Attention Is All You Need Paper</strong></a></li>\n</ul>\n<p><br></p>\n<h3><b>Model Details</b></h3>\n<ul>\n<li><strong>Encoder</strong>: <ul>\n<li><strong><em>EfficientNetV2-B2</em></strong> - Weights From <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\"><strong>Here</strong></a></li></ul></li>\n<li><strong>Decoder:</strong> <ul>\n<li><em>Transformer</em></li>\n<li><em>Number of Layers -&gt; 4</em></li>\n<li><em>Number of Heads -&gt; 8</em></li>\n<li><strong><em><code>d_model</code></em></strong> <em>-&gt; Image Embedding Depth (<strong><code>208</code></strong>)</em></li>\n<li><em>Units In Pointwise Feed-Forward NN -&gt; 512</em></li>\n<li><em>Dropout -&gt; 0.1</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Data Details</b></h3>\n<ul>\n<li><strong>Data Size:</strong><ul>\n<li><em>Val -&gt; 80,000 (random sample)</em></li>\n<li><em>Train -&gt; All-80,000</em></li>\n<li><em>Test -&gt; All</em></li></ul></li>\n<li><strong>Batch Size:</strong><ul>\n<li><em>128 Per Replica (1024 overall)</em></li></ul></li>\n<li><strong>Image Size:</strong><ul>\n<li><em>192x384x3</em></li></ul></li>\n<li><strong>Image Prep:</strong><ul>\n<li><em>TRAIN: Crop To Molecule, Invert and Resize (no padding)</em></li>\n<li><em>TEST : Crop To Molecule, Invert and Resize (no padding) … <strong>also if h&gt;w rotate 90 degrees</strong></em></li></ul></li>\n<li><strong>InChI Manipulation:</strong><ul>\n<li><em>Limit Max Length to Be Half Of Actual Max Length</em></li></ul></li>\n<li><strong>InChI Tokenization:</strong><ul>\n<li><em>Entities Are Considered Tokens (i.e. <code>Br</code> is a token, <code>191</code> is a token, etc.)</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Pipeline and General Details</b></h3>\n<ul>\n<li><strong>Execution &amp; Development Environment</strong><ul>\n<li><em>Kaggle Notebook</em></li></ul></li>\n<li><strong>Framework</strong><ul>\n<li><em>Tensorflow</em></li></ul></li>\n<li><strong>Accelerator</strong><ul>\n<li><em>TPU v3-8</em></li></ul></li>\n<li><strong>Data Format</strong><ul>\n<li><em>TFRecords</em></li></ul></li>\n<li><strong>I/O Streaming Tools</strong><ul>\n<li><em>tf.data</em></li></ul></li>\n<li><strong>Learning Rate Scheduler</strong><ul>\n<li><strong><em>Implemented from the \"Attention is All You Need Paper\" at half the magnitude… still too high. Very unstable.</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Result Details</b></h3>\n<ul>\n<li><strong>Timing</strong><ul>\n<li><em>~1.5-2.5 hours to complete training</em></li>\n<li><em>~8-12 minutes per epoch (includes validation)</em></li>\n<li><em>~40-45 minutes for the entire inference and submission process</em></li>\n<li><em>~35 minutes to infer on the test data</em></li>\n<li><em>8 minutes and 32 seconds to post-process test data (tokens to strings)</em></li>\n<li><strong><em>~2.50-4.25 hours to run the entire notebook including extras and visualizations</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<ul>\n<li><strong>Validation Levenshtein Distance (up to half max length) By Epoch (every other)</strong><ul>\n<li><em>2  -&gt; 11.55 CV LSD</em></li>\n<li><em>4  -&gt; 4.85 CV LSD</em></li>\n<li><strong><em>6  -&gt; 3.71 CV LSD</em></strong></li></ul></li>\n</ul>\n<p><br></p>\n<ul>\n<li><strong>Final Validation Levenshtein Distance</strong><ul>\n<li><strong><em>NOT COMPUTED</em></strong></li></ul></li>\n<li><strong>Public Leaderboard Score</strong><ul>\n<li><strong><em>5.83 LB LSD</em></strong></li></ul></li>\n</ul>\n<p>This was just where I had to stop because I ran out of TPU… I assume with a better schedule you could take this a lot further. Even on half-length tokens.</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p>I hope this helps! It was an effort and a half to get this to work and I'm certain that I did things that don't make sense or could be better. If you see anything like this please don't hesitate to let me know. I'm all about learning and getting better.</p>\n<p><br></p>\n<hr>\n<p>Thank you for your time!</p>",
      "rawMarkdown": "Hi there, \n\nI have created a [notebook](https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e) that implements a pipeline consisting of a Transformer Decoder and an EfficientNetV2 Encoder (all in Tensorflow).\n\nThis was a bit more difficult than I anticipated and I learned a lot (it took all 30 TPU hours this week and I still didn't get inference on the test set working). I just want to take this opportunity and share the knowledge and struggles with everyone.\n\nThe full architecture was able to achieve **CV=3.71 in 6 epochs** *(although I loaded the EfficientNetV2 pre-trained model from my other pipeline as a starting point for the encoder)*\n\nThe learning rate scheduling is difficult to get right and I encountered lots and lots of trouble with spiking loss and weird stuff happening. I think batching by sequence length and using a lower max LR would help considerably.\n\n---\n\n<br>\n\n***TL;DR  –––––––  TF+TPU+EfficientNetV2+Decoder  –––>* DIFFICULT TO TRAIN ... BUT 🔥🔥🧑‍🚒(me)🔥🔥**\n\n<br>\n\n---\n\n<br>\n\n*A **caveat**: I ran out of TPU hours and was not able to get the test inference to work. Also, the loss/accuracy became unstable during my last training (LR curve should be flatter and lower probably). All this being equal, the validation inference DOES work, and it's very similar to the test inference... so I must just be missing something. But as such, I have not submitted this notebook to the public LB yet. If someone gets it to work please let me know how it does!*\n\n***ADDENDUM***: \n* Apparently my notebook continued after my TPU ran out. I just saw that my notebook was able to START inferring on the test set (who knows if it will finish and if I can save it). It will take about 1300 seconds to infer across the whole test set and another 5-10 for post-processing. \n* So ~30 minutes for total inference time on the test dataset using this pipeline\n\n***FINAL ADDENDUM***: \n* Somehow my save and submission went through. This notebook scores **5.83 LB LSD**\n<br>\n\n---\n\n<br>\n\n### <b>Links</b>\n* [**My End-to-End Transformer+EfficientNetV2 Pipeline Notebook**](https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e)\n* [**My Previous EfficientNetV2+LSTM End-to-End Pipeline Notebook**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)\n* [**My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2**](https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2)\n* [**How To Load ImageNet Weights Into EfficientNetV2 And Infer** ***(for fine-tuning)***](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights)\n* [**The EfficientNetV2 Paper**](https://arxiv.org/pdf/2104.00298.pdf)\n* [**The Official Google Implementation of EfficientNetV2**](https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2)\n* [**The Show, Attend, and Tell Paper**](https://arxiv.org/pdf/1502.03044)\n* [**The Attention Is All You Need Paper**](https://arxiv.org/abs/1706.03762)\n\n<br>\n\n### <b>Model Details</b>\n\n* **Encoder**: \n  * ***EfficientNetV2-B2** - Weights From [**Here**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)*\n* **Decoder:** \n  * *Transformer*\n    * *Number of Layers -> 4*\n    * *Number of Heads -> 8*\n    * ***`d_model`*** *-> Image Embedding Depth (**`208`**)*\n    * *Units In Pointwise Feed-Forward NN -> 512*\n    * *Dropout -> 0.1*\n\n<br>\n\n### <b>Data Details</b>\n\n* **Data Size:**\n  * *Val -> 80,000 (random sample)*\n  * *Train -> All-80,000*\n  * *Test -> All*\n* **Batch Size:**\n  * *128 Per Replica (1024 overall)*\n* **Image Size:**\n  * *192x384x3*\n* **Image Prep:**\n  * *TRAIN: Crop To Molecule, Invert and Resize (no padding)*\n  * *TEST : Crop To Molecule, Invert and Resize (no padding) ... **also if h>w rotate 90 degrees***\n* **InChI Manipulation:**\n  * *Limit Max Length to Be Half Of Actual Max Length*\n* **InChI Tokenization:**\n  * *Entities Are Considered Tokens (i.e. `Br` is a token, `191` is a token, etc.)*\n\n<br>\n\n### <b>Pipeline and General Details</b>\n\n* **Execution & Development Environment**\n  * *Kaggle Notebook*\n* **Framework**\n  * *Tensorflow*\n* **Accelerator**\n  * *TPU v3-8*\n* **Data Format**\n  * *TFRecords*\n* **I/O Streaming Tools**\n  * *tf.data*\n* **Learning Rate Scheduler**\n  * ***Implemented from the \"Attention is All You Need Paper\" at half the magnitude... still too high. Very unstable.***\n\n<br>\n\n### <b>Result Details</b>\n \n* **Timing**\n  * *~1.5-2.5 hours to complete training*\n    * *~8-12 minutes per epoch (includes validation)*\n  * *~40-45 minutes for the entire inference and submission process*\n    * *~35 minutes to infer on the test data*\n    * *8 minutes and 32 seconds to post-process test data (tokens to strings)*\n  * ***~2.50-4.25 hours to run the entire notebook including extras and visualizations***\n\n<br>\n\n* **Validation Levenshtein Distance (up to half max length) By Epoch (every other)**\n  * *2  -> 11.55 CV LSD*\n  * *4  -> 4.85 CV LSD*\n  * ***6  -> 3.71 CV LSD***\n\n<br>\n\n* **Final Validation Levenshtein Distance**\n  * ***NOT COMPUTED***\n* **Public Leaderboard Score**\n  * ***5.83 LB LSD***\n\nThis was just where I had to stop because I ran out of TPU... I assume with a better schedule you could take this a lot further. Even on half-length tokens.\n\n<br>\n\n---\n\n<br>\n\nI hope this helps! It was an effort and a half to get this to work and I'm certain that I did things that don't make sense or could be better. If you see anything like this please don't hesitate to let me know. I'm all about learning and getting better.\n\n<br>\n\n---\n\nThank you for your time!",
      "votes": null
    },
    {
      "id": "1323453",
      "postDate": "05/26/2021 08:37:06",
      "content": "<p>I'm an extreme begginer here, but for training of transformers , DeepSpeed and DDP can be used to significantly amp up the training speed.  Also I dont think taking RGB values helps in making the algorithm learn something extra and only grayscale pics can help you reduces a bit of computaion time. Please let me know if i was of any help, it would be helpful for me!</p>",
      "rawMarkdown": "I'm an extreme begginer here, but for training of transformers , DeepSpeed and DDP can be used to significantly amp up the training speed.  Also I dont think taking RGB values helps in making the algorithm learn something extra and only grayscale pics can help you reduces a bit of computaion time. Please let me know if i was of any help, it would be helpful for me!",
      "votes": null
    },
    {
      "id": "1325030",
      "postDate": "05/27/2021 12:55:46",
      "content": "<p>how could you find spike learning rate is key point,</p>\n<ol>\n<li>In my intuitive if you  training more epoch every loss will begin to decrease ( untill this time  )</li>\n<li>at first I think the problem is model capacity is not enough , I add more layer or unit , but give very little improve </li>\n<li>then I try to use pre-training weight but still seems not work</li>\n<li>I even think is bfloat16 weak language type convert cause unstable</li>\n</ol>\n<p>I read \"attention is all you need\" paper but fully ignore this part<br>\nI waste two weeks time on dig in model structure to find something went wrong<br>\nThank you for reply in my post  to point this part out </p>",
      "rawMarkdown": "how could you find spike learning rate is key point,\n\n1. In my intuitive if you  training more epoch every loss will begin to decrease ( untill this time  )\n2. at first I think the problem is model capacity is not enough , I add more layer or unit , but give very little improve \n3. then I try to use pre-training weight but still seems not work\n4. I even think is bfloat16 weak language type convert cause unstable\n\nI read \"attention is all you need\" paper but fully ignore this part\nI waste two weeks time on dig in model structure to find something went wrong\nThank you for reply in my post  to point this part out",
      "votes": null
    },
    {
      "id": "1325060",
      "postDate": "05/27/2021 13:20:47",
      "content": "<p>I think it's the learning rate is too high. It also might be due to the image captioning flow?</p>\n<p>When I tried freezing the encoder (using a pre-trained one), I noticed it was much more stable.</p>\n<p>Also, the <strong><code>d_model</code></strong> size is very small in the above experiment. This might make it unstable. If you use an EfficientNetV1 model the d_model variable will be 5-10 times larger and the Transformer will get waayyyy bigger. Maybe this is impacting it too?</p>\n<p>When the competition is over I will opensource my private notebook that does training using a larger image model and a slightly modified transformer from above.</p>\n<p>Glad it helped!</p>",
      "rawMarkdown": "I think it's the learning rate is too high. It also might be due to the image captioning flow?\n\nWhen I tried freezing the encoder (using a pre-trained one), I noticed it was much more stable.\n\nAlso, the **`d_model`** size is very small in the above experiment. This might make it unstable. If you use an EfficientNetV1 model the d_model variable will be 5-10 times larger and the Transformer will get waayyyy bigger. Maybe this is impacting it too?\n\nWhen the competition is over I will opensource my private notebook that does training using a larger image model and a slightly modified transformer from above.\n\nGlad it helped!",
      "votes": null
    },
    {
      "id": "1325063",
      "postDate": "05/27/2021 13:25:52",
      "content": "<p>RGB does not help over grayscale, but it doesn't really increase the network size that much and it's easier to implement. Print a network summary using one channel vs 3-channel and you'll see the difference is minimal.</p>\n<p>As for DeepSpeed and DDP, I haven't heard or read about this. Can you link me to anything specific? I can try to implement it?</p>\n<p>Thanks for your comment!</p>",
      "rawMarkdown": "RGB does not help over grayscale, but it doesn't really increase the network size that much and it's easier to implement. Print a network summary using one channel vs 3-channel and you'll see the difference is minimal.\n\nAs for DeepSpeed and DDP, I haven't heard or read about this. Can you link me to anything specific? I can try to implement it?\n\nThanks for your comment!",
      "votes": null
    },
    {
      "id": "1325199",
      "postDate": "05/27/2021 15:08:17",
      "content": "<p>Deepspeed is Microsoft’s library for training very large distributed models in various ways and implements zero redundancy optimizers and some other fancy things. It’s pytorch only though and won’t be of much use here. DDP is just pytorch native distributed data parallel. </p>",
      "rawMarkdown": "Deepspeed is Microsoft’s library for training very large distributed models in various ways and implements zero redundancy optimizers and some other fancy things. It’s pytorch only though and won’t be of much use here. DDP is just pytorch native distributed data parallel.",
      "votes": null
    },
    {
      "id": "1325271",
      "postDate": "05/27/2021 15:59:34",
      "content": "<p>TF2.X already handles distributed training very efficiently with almost no code change. This wasn't the case with TF1.X</p>\n<p>Setting up DDP with Pytorch wasn't an easy experience for me (I didn't use Lightning though but vanilla Pytorch instead)</p>",
      "rawMarkdown": "TF2.X already handles distributed training very efficiently with almost no code change. This wasn't the case with TF1.X\n\nSetting up DDP with Pytorch wasn't an easy experience for me (I didn't use Lightning though but vanilla Pytorch instead)",
      "votes": null
    },
    {
      "id": "1325801",
      "postDate": "05/28/2021 02:45:26",
      "content": "<p>Thanks a lot for the great share. Any suggestion for how to match LR if I start train from previously trained weight for lower max LR? </p>",
      "rawMarkdown": "Thanks a lot for the great share. Any suggestion for how to match LR if I start train from previously trained weight for lower max LR?",
      "votes": null
    },
    {
      "id": "1325848",
      "postDate": "05/28/2021 04:17:17",
      "content": "<p>why does transformer better than LSTM in precise?<br>\nDoes you try vision transformer? I plan to try it asap</p>",
      "rawMarkdown": "why does transformer better than LSTM in precise?\nDoes you try vision transformer? I plan to try it asap",
      "votes": null
    },
    {
      "id": "1326422",
      "postDate": "05/28/2021 13:30:15",
      "content": "<p>Thanks a lot for sharing.</p>",
      "rawMarkdown": "Thanks a lot for sharing.",
      "votes": null
    },
    {
      "id": "1326759",
      "postDate": "05/28/2021 16:59:54",
      "content": "<p>I'm struggling with this now. I think because I am constrained to Kaggle's limits, I lose some information when restarting training. My only suggestion would be to lower the peak LR for each subsequent training session. i.e. increase the value in the rsqrt().</p>",
      "rawMarkdown": "I'm struggling with this now. I think because I am constrained to Kaggle's limits, I lose some information when restarting training. My only suggestion would be to lower the peak LR for each subsequent training session. i.e. increase the value in the rsqrt().",
      "votes": null
    },
    {
      "id": "1331067",
      "postDate": "06/01/2021 08:53:46",
      "content": "<p>Validation take about 30-40 minutes? am I config something wrong?</p>",
      "rawMarkdown": "Validation take about 30-40 minutes? am I config something wrong?",
      "votes": null
    },
    {
      "id": "1331418",
      "postDate": "06/01/2021 13:08:10",
      "content": "<p>The first validation run seems to take longer than the following runs. 30-40 minutes seems excessive though. I found it only added a few extra minutes per epoch.</p>\n<p>That being said, the transformer will be inferring <strong><code>MAX_LEN</code></strong> times more frequently during validation/testing. So some increase in timing is expected.</p>",
      "rawMarkdown": "The first validation run seems to take longer than the following runs. 30-40 minutes seems excessive though. I found it only added a few extra minutes per epoch.\n\nThat being said, the transformer will be inferring **`MAX_LEN`** times more frequently during validation/testing. So some increase in timing is expected.",
      "votes": null
    },
    {
      "id": "1332562",
      "postDate": "06/02/2021 07:25:07",
      "content": "<p>it seem valid step doesn't use TPU , strangely</p>",
      "rawMarkdown": "it seem valid step doesn't use TPU , strangely",
      "votes": null
    },
    {
      "id": "1332787",
      "postDate": "06/02/2021 09:52:24",
      "content": "<p>A great notebook - can be a good guidance for Image Processing using TPUs. As I understand, the image processing is done separately (couldn't see the processing of the original /train and /test images) - how long did that part take?</p>",
      "rawMarkdown": "A great notebook - can be a good guidance for Image Processing using TPUs. As I understand, the image processing is done separately (couldn't see the processing of the original /train and /test images) - how long did that part take?",
      "votes": null
    },
    {
      "id": "1333495",
      "postDate": "06/02/2021 18:58:37",
      "content": "<p>Thanks for sharing your progress throughout the competition! I haven't coded for TPU yet but would like to compare models. My implementation (on GPU) takes MUCH longer than your TPU implementation. Do you know what your training time would look like on GPU?</p>",
      "rawMarkdown": "Thanks for sharing your progress throughout the competition! I haven't coded for TPU yet but would like to compare models. My implementation (on GPU) takes MUCH longer than your TPU implementation. Do you know what your training time would look like on GPU?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1323453,
      "author_name": "puranjaymishra",
      "author_url": "",
      "post_date": "05/26/2021 08:37:06",
      "content": "<p>I'm an extreme begginer here, but for training of transformers , DeepSpeed and DDP can be used to significantly amp up the training speed.  Also I dont think taking RGB values helps in making the algorithm learn something extra and only grayscale pics can help you reduces a bit of computaion time. Please let me know if i was of any help, it would be helpful for me!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1325063,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/27/2021 13:25:52",
          "content": "<p>RGB does not help over grayscale, but it doesn't really increase the network size that much and it's easier to implement. Print a network summary using one channel vs 3-channel and you'll see the difference is minimal.</p>\n<p>As for DeepSpeed and DDP, I haven't heard or read about this. Can you link me to anything specific? I can try to implement it?</p>\n<p>Thanks for your comment!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1325199,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "05/27/2021 15:08:17",
          "content": "<p>Deepspeed is Microsoft’s library for training very large distributed models in various ways and implements zero redundancy optimizers and some other fancy things. It’s pytorch only though and won’t be of much use here. DDP is just pytorch native distributed data parallel. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1325271,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/27/2021 15:59:34",
          "content": "<p>TF2.X already handles distributed training very efficiently with almost no code change. This wasn't the case with TF1.X</p>\n<p>Setting up DDP with Pytorch wasn't an easy experience for me (I didn't use Lightning though but vanilla Pytorch instead)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1325030,
      "author_name": "drzhuzhe",
      "author_url": "",
      "post_date": "05/27/2021 12:55:46",
      "content": "<p>how could you find spike learning rate is key point,</p>\n<ol>\n<li>In my intuitive if you  training more epoch every loss will begin to decrease ( untill this time  )</li>\n<li>at first I think the problem is model capacity is not enough , I add more layer or unit , but give very little improve </li>\n<li>then I try to use pre-training weight but still seems not work</li>\n<li>I even think is bfloat16 weak language type convert cause unstable</li>\n</ol>\n<p>I read \"attention is all you need\" paper but fully ignore this part<br>\nI waste two weeks time on dig in model structure to find something went wrong<br>\nThank you for reply in my post  to point this part out </p>",
      "votes": null,
      "replies": [
        {
          "id": 1325060,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/27/2021 13:20:47",
          "content": "<p>I think it's the learning rate is too high. It also might be due to the image captioning flow?</p>\n<p>When I tried freezing the encoder (using a pre-trained one), I noticed it was much more stable.</p>\n<p>Also, the <strong><code>d_model</code></strong> size is very small in the above experiment. This might make it unstable. If you use an EfficientNetV1 model the d_model variable will be 5-10 times larger and the Transformer will get waayyyy bigger. Maybe this is impacting it too?</p>\n<p>When the competition is over I will opensource my private notebook that does training using a larger image model and a slightly modified transformer from above.</p>\n<p>Glad it helped!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1325848,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "05/28/2021 04:17:17",
          "content": "<p>why does transformer better than LSTM in precise?<br>\nDoes you try vision transformer? I plan to try it asap</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1325801,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "05/28/2021 02:45:26",
      "content": "<p>Thanks a lot for the great share. Any suggestion for how to match LR if I start train from previously trained weight for lower max LR? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1326759,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/28/2021 16:59:54",
          "content": "<p>I'm struggling with this now. I think because I am constrained to Kaggle's limits, I lose some information when restarting training. My only suggestion would be to lower the peak LR for each subsequent training session. i.e. increase the value in the rsqrt().</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1326422,
      "author_name": "saivineetha1998",
      "author_url": "",
      "post_date": "05/28/2021 13:30:15",
      "content": "<p>Thanks a lot for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1331067,
      "author_name": "drzhuzhe",
      "author_url": "",
      "post_date": "06/01/2021 08:53:46",
      "content": "<p>Validation take about 30-40 minutes? am I config something wrong?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1331418,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "06/01/2021 13:08:10",
          "content": "<p>The first validation run seems to take longer than the following runs. 30-40 minutes seems excessive though. I found it only added a few extra minutes per epoch.</p>\n<p>That being said, the transformer will be inferring <strong><code>MAX_LEN</code></strong> times more frequently during validation/testing. So some increase in timing is expected.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332562,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "06/02/2021 07:25:07",
          "content": "<p>it seem valid step doesn't use TPU , strangely</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332787,
      "author_name": "jhasanov",
      "author_url": "",
      "post_date": "06/02/2021 09:52:24",
      "content": "<p>A great notebook - can be a good guidance for Image Processing using TPUs. As I understand, the image processing is done separately (couldn't see the processing of the original /train and /test images) - how long did that part take?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333495,
      "author_name": "mvenou",
      "author_url": "",
      "post_date": "06/02/2021 18:58:37",
      "content": "<p>Thanks for sharing your progress throughout the competition! I haven't coded for TPU yet but would like to compare models. My implementation (on GPU) takes MUCH longer than your TPU implementation. Do you know what your training time would look like on GPU?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1322932": "Hi there, \n\nI have created a [notebook](https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e) that implements a pipeline consisting of a Transformer Decoder and an EfficientNetV2 Encoder (all in Tensorflow).\n\nThis was a bit more difficult than I anticipated and I learned a lot (it took all 30 TPU hours this week and I still didn't get inference on the test set working). I just want to take this opportunity and share the knowledge and struggles with everyone.\n\nThe full architecture was able to achieve **CV=3.71 in 6 epochs** *(although I loaded the EfficientNetV2 pre-trained model from my other pipeline as a starting point for the encoder)*\n\nThe learning rate scheduling is difficult to get right and I encountered lots and lots of trouble with spiking loss and weird stuff happening. I think batching by sequence length and using a lower max LR would help considerably.\n\n---\n\n<br>\n\n***TL;DR  –––––––  TF+TPU+EfficientNetV2+Decoder  –––>* DIFFICULT TO TRAIN ... BUT 🔥🔥🧑‍🚒(me)🔥🔥**\n\n<br>\n\n---\n\n<br>\n\n*A **caveat**: I ran out of TPU hours and was not able to get the test inference to work. Also, the loss/accuracy became unstable during my last training (LR curve should be flatter and lower probably). All this being equal, the validation inference DOES work, and it's very similar to the test inference... so I must just be missing something. But as such, I have not submitted this notebook to the public LB yet. If someone gets it to work please let me know how it does!*\n\n***ADDENDUM***: \n* Apparently my notebook continued after my TPU ran out. I just saw that my notebook was able to START inferring on the test set (who knows if it will finish and if I can save it). It will take about 1300 seconds to infer across the whole test set and another 5-10 for post-processing. \n* So ~30 minutes for total inference time on the test dataset using this pipeline\n\n***FINAL ADDENDUM***: \n* Somehow my save and submission went through. This notebook scores **5.83 LB LSD**\n<br>\n\n---\n\n<br>\n\n### <b>Links</b>\n* [**My End-to-End Transformer+EfficientNetV2 Pipeline Notebook**](https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e)\n* [**My Previous EfficientNetV2+LSTM End-to-End Pipeline Notebook**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)\n* [**My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2**](https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2)\n* [**How To Load ImageNet Weights Into EfficientNetV2 And Infer** ***(for fine-tuning)***](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights)\n* [**The EfficientNetV2 Paper**](https://arxiv.org/pdf/2104.00298.pdf)\n* [**The Official Google Implementation of EfficientNetV2**](https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2)\n* [**The Show, Attend, and Tell Paper**](https://arxiv.org/pdf/1502.03044)\n* [**The Attention Is All You Need Paper**](https://arxiv.org/abs/1706.03762)\n\n<br>\n\n### <b>Model Details</b>\n\n* **Encoder**: \n  * ***EfficientNetV2-B2** - Weights From [**Here**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)*\n* **Decoder:** \n  * *Transformer*\n    * *Number of Layers -> 4*\n    * *Number of Heads -> 8*\n    * ***`d_model`*** *-> Image Embedding Depth (**`208`**)*\n    * *Units In Pointwise Feed-Forward NN -> 512*\n    * *Dropout -> 0.1*\n\n<br>\n\n### <b>Data Details</b>\n\n* **Data Size:**\n  * *Val -> 80,000 (random sample)*\n  * *Train -> All-80,000*\n  * *Test -> All*\n* **Batch Size:**\n  * *128 Per Replica (1024 overall)*\n* **Image Size:**\n  * *192x384x3*\n* **Image Prep:**\n  * *TRAIN: Crop To Molecule, Invert and Resize (no padding)*\n  * *TEST : Crop To Molecule, Invert and Resize (no padding) ... **also if h>w rotate 90 degrees***\n* **InChI Manipulation:**\n  * *Limit Max Length to Be Half Of Actual Max Length*\n* **InChI Tokenization:**\n  * *Entities Are Considered Tokens (i.e. `Br` is a token, `191` is a token, etc.)*\n\n<br>\n\n### <b>Pipeline and General Details</b>\n\n* **Execution & Development Environment**\n  * *Kaggle Notebook*\n* **Framework**\n  * *Tensorflow*\n* **Accelerator**\n  * *TPU v3-8*\n* **Data Format**\n  * *TFRecords*\n* **I/O Streaming Tools**\n  * *tf.data*\n* **Learning Rate Scheduler**\n  * ***Implemented from the \"Attention is All You Need Paper\" at half the magnitude... still too high. Very unstable.***\n\n<br>\n\n### <b>Result Details</b>\n \n* **Timing**\n  * *~1.5-2.5 hours to complete training*\n    * *~8-12 minutes per epoch (includes validation)*\n  * *~40-45 minutes for the entire inference and submission process*\n    * *~35 minutes to infer on the test data*\n    * *8 minutes and 32 seconds to post-process test data (tokens to strings)*\n  * ***~2.50-4.25 hours to run the entire notebook including extras and visualizations***\n\n<br>\n\n* **Validation Levenshtein Distance (up to half max length) By Epoch (every other)**\n  * *2  -> 11.55 CV LSD*\n  * *4  -> 4.85 CV LSD*\n  * ***6  -> 3.71 CV LSD***\n\n<br>\n\n* **Final Validation Levenshtein Distance**\n  * ***NOT COMPUTED***\n* **Public Leaderboard Score**\n  * ***5.83 LB LSD***\n\nThis was just where I had to stop because I ran out of TPU... I assume with a better schedule you could take this a lot further. Even on half-length tokens.\n\n<br>\n\n---\n\n<br>\n\nI hope this helps! It was an effort and a half to get this to work and I'm certain that I did things that don't make sense or could be better. If you see anything like this please don't hesitate to let me know. I'm all about learning and getting better.\n\n<br>\n\n---\n\nThank you for your time!",
    "1323453": "I'm an extreme begginer here, but for training of transformers , DeepSpeed and DDP can be used to significantly amp up the training speed.  Also I dont think taking RGB values helps in making the algorithm learn something extra and only grayscale pics can help you reduces a bit of computaion time. Please let me know if i was of any help, it would be helpful for me!",
    "1325030": "how could you find spike learning rate is key point,\n\n1. In my intuitive if you  training more epoch every loss will begin to decrease ( untill this time  )\n2. at first I think the problem is model capacity is not enough , I add more layer or unit , but give very little improve \n3. then I try to use pre-training weight but still seems not work\n4. I even think is bfloat16 weak language type convert cause unstable\n\nI read \"attention is all you need\" paper but fully ignore this part\nI waste two weeks time on dig in model structure to find something went wrong\nThank you for reply in my post  to point this part out",
    "1325060": "I think it's the learning rate is too high. It also might be due to the image captioning flow?\n\nWhen I tried freezing the encoder (using a pre-trained one), I noticed it was much more stable.\n\nAlso, the **`d_model`** size is very small in the above experiment. This might make it unstable. If you use an EfficientNetV1 model the d_model variable will be 5-10 times larger and the Transformer will get waayyyy bigger. Maybe this is impacting it too?\n\nWhen the competition is over I will opensource my private notebook that does training using a larger image model and a slightly modified transformer from above.\n\nGlad it helped!",
    "1325063": "RGB does not help over grayscale, but it doesn't really increase the network size that much and it's easier to implement. Print a network summary using one channel vs 3-channel and you'll see the difference is minimal.\n\nAs for DeepSpeed and DDP, I haven't heard or read about this. Can you link me to anything specific? I can try to implement it?\n\nThanks for your comment!",
    "1325199": "Deepspeed is Microsoft’s library for training very large distributed models in various ways and implements zero redundancy optimizers and some other fancy things. It’s pytorch only though and won’t be of much use here. DDP is just pytorch native distributed data parallel.",
    "1325271": "TF2.X already handles distributed training very efficiently with almost no code change. This wasn't the case with TF1.X\n\nSetting up DDP with Pytorch wasn't an easy experience for me (I didn't use Lightning though but vanilla Pytorch instead)",
    "1325801": "Thanks a lot for the great share. Any suggestion for how to match LR if I start train from previously trained weight for lower max LR?",
    "1325848": "why does transformer better than LSTM in precise?\nDoes you try vision transformer? I plan to try it asap",
    "1326422": "Thanks a lot for sharing.",
    "1326759": "I'm struggling with this now. I think because I am constrained to Kaggle's limits, I lose some information when restarting training. My only suggestion would be to lower the peak LR for each subsequent training session. i.e. increase the value in the rsqrt().",
    "1331067": "Validation take about 30-40 minutes? am I config something wrong?",
    "1331418": "The first validation run seems to take longer than the following runs. 30-40 minutes seems excessive though. I found it only added a few extra minutes per epoch.\n\nThat being said, the transformer will be inferring **`MAX_LEN`** times more frequently during validation/testing. So some increase in timing is expected.",
    "1332562": "it seem valid step doesn't use TPU , strangely",
    "1332787": "A great notebook - can be a good guidance for Image Processing using TPUs. As I understand, the image processing is done separately (couldn't see the processing of the original /train and /test images) - how long did that part take?",
    "1333495": "Thanks for sharing your progress throughout the competition! I haven't coded for TPU yet but would like to compare models. My implementation (on GPU) takes MUCH longer than your TPU implementation. Do you know what your training time would look like on GPU?"
  },
  "source": "meta"
}