{
  "id": 239595,
  "title": "EfficientNetV2 Encoder Implemented [4.97 LB] – Training+Inference In Under 3 Hours 🔥🔥🔥",
  "url": "/competitions/bms-molecular-translation/discussion/239595",
  "author_name": "",
  "post_date": "2021-05-16T21:00:28.125339500Z",
  "votes": 46,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi there, </p>\n<p>I have created a notebook that implements the EfficientNetV2 model architecture (in Tensorflow) to function as the image encoder.</p>\n<p>Not only that. I implement it in such a way that <strong>BOTH THE TRAINING AND THE MODEL INFERENCE ON THE TEST SET</strong> can happen in <strong>APPROXIMATELY 3 HOURS (OR LESS!)</strong> </p>\n<hr>\n<p><br></p>\n<p><strong><em>TL;DR  –––––––  TF+TPU+EfficientNetV2  –––&gt;</em> 🔥🔥🔥🔥🔥🔥🔥🔥</strong></p>\n<p><br></p>\n<hr>\n<p><em>A caveat to the work I've completed. As this notebook implements both of these things, I don't have a comparative version with EfficientNetB2 to benchmark against. If I have time in the future I will do this just so we can all see the comparative performance differences.</em></p>\n<p><em>A second caveat is that I did the notebook development and initial submission within an interactive session. Therefore, there MAY be problems with a save-run-all scenario… I hope not. If I have TPU near the end of the week I will re-run this notebook using the run-all save option. If someone else does this and it succeeds with no alterations please let me know. Thanks</em> 😊</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h3><b>Links</b></h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/edit\" target=\"_blank\"><strong>My End-to-End Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2\" target=\"_blank\"><strong>My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\"><strong>How To Load ImageNet Weights Into EfficientNetV2 And Infer</strong> <strong><em>(for fine-tuning)</em></strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\"><strong>The EfficientNetV2 Paper</strong></a></li>\n<li><a href=\"https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2\" target=\"_blank\"><strong>The Official Google Implementation of EfficientNetV2</strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/1502.03044\" target=\"_blank\"><strong>The Show, Attend, and Tell Paper</strong></a></li>\n</ul>\n<p><br></p>\n<h3><b>Model Details</b></h3>\n<ul>\n<li><strong>Encoder</strong>: <ul>\n<li><strong><em>EfficientNetV2-B2</em></strong> (no pre-trained weights loaded)</li></ul></li>\n<li><strong>Attention:</strong><ul>\n<li><em><a href=\"https://arxiv.org/abs/1409.0473\" target=\"_blank\">Bhadanau</a> w/ 192 Units</em></li></ul></li>\n<li><strong>Decoder:</strong> <ul>\n<li><em>LSTM w/ 512 Units and 0.05 Dropout</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Data Details</b></h3>\n<ul>\n<li><strong>Data Size:</strong><ul>\n<li><em>Val -&gt; 80,000 (random sample)</em></li>\n<li><em>Train -&gt; All-80,000</em></li>\n<li><em>Test -&gt; All</em></li></ul></li>\n<li><strong>Batch Size:</strong><ul>\n<li><em>64 Per Replica (512 overall) … this could have been higher in hindsight (128/1024)</em></li></ul></li>\n<li><strong>Image Size:</strong><ul>\n<li><em>192x384x3</em></li></ul></li>\n<li><strong>Image Prep:</strong><ul>\n<li><em>TRAIN: Crop To Molecule, Invert and Resize (no padding)</em></li>\n<li><em>TEST : Crop To Molecule, Invert and Resize (no padding) … <strong>also if h&gt;w rotate 90 degrees</strong></em></li></ul></li>\n<li><strong>InChI Manipulation:</strong><ul>\n<li><em>Limit Max Length to Be Half Of Actual Max Length</em></li></ul></li>\n<li><strong>InChI Tokenization:</strong><ul>\n<li><em>Entities Are Considered Tokens (i.e. <code>Br</code> is a token, <code>191</code> is a token, etc.)</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Pipeline and General Details</b></h3>\n<ul>\n<li><strong>Execution &amp; Development Environment</strong><ul>\n<li><em>Kaggle Notebook</em></li></ul></li>\n<li><strong>Framework</strong><ul>\n<li><em>Tensorflow</em></li></ul></li>\n<li><strong>Accelerator</strong><ul>\n<li><em>TPU v3-8</em></li></ul></li>\n<li><strong>Data Format</strong><ul>\n<li><em>TFRecords</em></li></ul></li>\n<li><strong>I/O Streaming Tools</strong><ul>\n<li><em>tf.data</em></li></ul></li>\n<li><strong>Learning Rate Scheduler</strong><ul>\n<li><em>Ramp up to <strong>2e-3</strong>… step down to <strong>1e-5</strong></em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Result Details</b></h3>\n<ul>\n<li><p><strong>Timing</strong></p>\n<ul>\n<li><em>~2-2.5 hours to complete training</em></li>\n<li><em>~10-12 minutes per epoch (includes validation)</em></li>\n<li><em>~53 seconds to infer on validation dataset afterwards for visualization's sake…</em></li>\n<li><em>14 minutes and 15 seconds for the entire inference and submission process</em></li>\n<li><em>5 minutes and 40 seconds to infer on test data</em></li>\n<li><em>8 minutes and 32 seconds to post-process test data (tokens to strings)</em></li>\n<li><strong><em>~2.75-3.25 hours to run the entire notebook including extras and visualizations</em></strong></li>\n<li><strong><em>~2.5-2.75 hours of required computation… the rest is for show or display purposes</em></strong></li></ul></li>\n<li><p><strong>Validation Levenshtein Distance (up to half max length) By Epoch</strong></p>\n<ul>\n<li><em>1  -&gt; 15.79</em></li>\n<li><em>2  -&gt; 9.51</em></li>\n<li><em>3  -&gt; 6.99</em></li>\n<li><em>4  -&gt; 6.74</em></li>\n<li><em>5  -&gt; 4.91</em></li>\n<li><em>6  -&gt; 4.88</em></li>\n<li><em>7  -&gt; 4.02</em></li>\n<li><em>8  -&gt; 4.01</em></li>\n<li><em>9  -&gt; 3.64</em> </li>\n<li><em>10 -&gt; 3.77</em></li>\n<li><em>11 -&gt; 3.42</em></li>\n<li><em>12 -&gt; 3.32</em></li></ul></li>\n<li><p><strong>Final Validation Levenshtein Distance</strong></p>\n<ul>\n<li><strong><em>3.75</em></strong> <em>(not sure why this is different than LSD after epoch 12…)</em></li></ul></li>\n<li><p><strong>Public Leaderboard Score</strong></p>\n<ul>\n<li><strong><em>4.97</em></strong></li></ul></li>\n</ul>\n<p><em>SIDE NOTE: If you fine-tune on the training dataset for another 3-5 epochs with the max length being the correct value, I would estimate an improvement of 0.75-1.5 LSD on the LB… just a guestimate from previous experience though…</em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong><em>NOTE:</em></strong></p>\n<p>The discrepancy between public LB LSD and validation LSD is to be expected for 2 reasons.</p>\n<ol>\n<li>The validation LSD is calculated assuming a max length equivalent to half of the actual max length</li>\n<li>The strategy in place currently to handle the rotation for test images is not sufficient. I apply the h&gt;w then rotate strategy… but it fails to capture many of the cases for which rotation is required. </li>\n</ol>\n<p><br></p>\n<hr>\n<p><br></p>\n<p>I hope this helps!</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong><em>p.s.</em></strong> <em>I will be updating the notebook tomorrow or early this week to include an option to train the EfficientNetV2 model from the provided pre-trained ImageNet weights.</em></p>\n<p><strong><em>p.p.s.</em></strong> <em>I will also tack on the RDKit normalization procedure to the end of this notebook as per <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>'s notebook. Not sure if it will help… but why not.</em></p>",
  "messages": [
    {
      "id": "1310687",
      "postDate": "05/16/2021 21:00:28",
      "content": "<p>Hi there, </p>\n<p>I have created a notebook that implements the EfficientNetV2 model architecture (in Tensorflow) to function as the image encoder.</p>\n<p>Not only that. I implement it in such a way that <strong>BOTH THE TRAINING AND THE MODEL INFERENCE ON THE TEST SET</strong> can happen in <strong>APPROXIMATELY 3 HOURS (OR LESS!)</strong> </p>\n<hr>\n<p><br></p>\n<p><strong><em>TL;DR  –––––––  TF+TPU+EfficientNetV2  –––&gt;</em> 🔥🔥🔥🔥🔥🔥🔥🔥</strong></p>\n<p><br></p>\n<hr>\n<p><em>A caveat to the work I've completed. As this notebook implements both of these things, I don't have a comparative version with EfficientNetB2 to benchmark against. If I have time in the future I will do this just so we can all see the comparative performance differences.</em></p>\n<p><em>A second caveat is that I did the notebook development and initial submission within an interactive session. Therefore, there MAY be problems with a save-run-all scenario… I hope not. If I have TPU near the end of the week I will re-run this notebook using the run-all save option. If someone else does this and it succeeds with no alterations please let me know. Thanks</em> 😊</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h3><b>Links</b></h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/edit\" target=\"_blank\"><strong>My End-to-End Pipeline Notebook</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2\" target=\"_blank\"><strong>My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\"><strong>How To Load ImageNet Weights Into EfficientNetV2 And Infer</strong> <strong><em>(for fine-tuning)</em></strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\"><strong>The EfficientNetV2 Paper</strong></a></li>\n<li><a href=\"https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2\" target=\"_blank\"><strong>The Official Google Implementation of EfficientNetV2</strong></a></li>\n<li><a href=\"https://arxiv.org/pdf/1502.03044\" target=\"_blank\"><strong>The Show, Attend, and Tell Paper</strong></a></li>\n</ul>\n<p><br></p>\n<h3><b>Model Details</b></h3>\n<ul>\n<li><strong>Encoder</strong>: <ul>\n<li><strong><em>EfficientNetV2-B2</em></strong> (no pre-trained weights loaded)</li></ul></li>\n<li><strong>Attention:</strong><ul>\n<li><em><a href=\"https://arxiv.org/abs/1409.0473\" target=\"_blank\">Bhadanau</a> w/ 192 Units</em></li></ul></li>\n<li><strong>Decoder:</strong> <ul>\n<li><em>LSTM w/ 512 Units and 0.05 Dropout</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Data Details</b></h3>\n<ul>\n<li><strong>Data Size:</strong><ul>\n<li><em>Val -&gt; 80,000 (random sample)</em></li>\n<li><em>Train -&gt; All-80,000</em></li>\n<li><em>Test -&gt; All</em></li></ul></li>\n<li><strong>Batch Size:</strong><ul>\n<li><em>64 Per Replica (512 overall) … this could have been higher in hindsight (128/1024)</em></li></ul></li>\n<li><strong>Image Size:</strong><ul>\n<li><em>192x384x3</em></li></ul></li>\n<li><strong>Image Prep:</strong><ul>\n<li><em>TRAIN: Crop To Molecule, Invert and Resize (no padding)</em></li>\n<li><em>TEST : Crop To Molecule, Invert and Resize (no padding) … <strong>also if h&gt;w rotate 90 degrees</strong></em></li></ul></li>\n<li><strong>InChI Manipulation:</strong><ul>\n<li><em>Limit Max Length to Be Half Of Actual Max Length</em></li></ul></li>\n<li><strong>InChI Tokenization:</strong><ul>\n<li><em>Entities Are Considered Tokens (i.e. <code>Br</code> is a token, <code>191</code> is a token, etc.)</em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Pipeline and General Details</b></h3>\n<ul>\n<li><strong>Execution &amp; Development Environment</strong><ul>\n<li><em>Kaggle Notebook</em></li></ul></li>\n<li><strong>Framework</strong><ul>\n<li><em>Tensorflow</em></li></ul></li>\n<li><strong>Accelerator</strong><ul>\n<li><em>TPU v3-8</em></li></ul></li>\n<li><strong>Data Format</strong><ul>\n<li><em>TFRecords</em></li></ul></li>\n<li><strong>I/O Streaming Tools</strong><ul>\n<li><em>tf.data</em></li></ul></li>\n<li><strong>Learning Rate Scheduler</strong><ul>\n<li><em>Ramp up to <strong>2e-3</strong>… step down to <strong>1e-5</strong></em></li></ul></li>\n</ul>\n<p><br></p>\n<h3><b>Result Details</b></h3>\n<ul>\n<li><p><strong>Timing</strong></p>\n<ul>\n<li><em>~2-2.5 hours to complete training</em></li>\n<li><em>~10-12 minutes per epoch (includes validation)</em></li>\n<li><em>~53 seconds to infer on validation dataset afterwards for visualization's sake…</em></li>\n<li><em>14 minutes and 15 seconds for the entire inference and submission process</em></li>\n<li><em>5 minutes and 40 seconds to infer on test data</em></li>\n<li><em>8 minutes and 32 seconds to post-process test data (tokens to strings)</em></li>\n<li><strong><em>~2.75-3.25 hours to run the entire notebook including extras and visualizations</em></strong></li>\n<li><strong><em>~2.5-2.75 hours of required computation… the rest is for show or display purposes</em></strong></li></ul></li>\n<li><p><strong>Validation Levenshtein Distance (up to half max length) By Epoch</strong></p>\n<ul>\n<li><em>1  -&gt; 15.79</em></li>\n<li><em>2  -&gt; 9.51</em></li>\n<li><em>3  -&gt; 6.99</em></li>\n<li><em>4  -&gt; 6.74</em></li>\n<li><em>5  -&gt; 4.91</em></li>\n<li><em>6  -&gt; 4.88</em></li>\n<li><em>7  -&gt; 4.02</em></li>\n<li><em>8  -&gt; 4.01</em></li>\n<li><em>9  -&gt; 3.64</em> </li>\n<li><em>10 -&gt; 3.77</em></li>\n<li><em>11 -&gt; 3.42</em></li>\n<li><em>12 -&gt; 3.32</em></li></ul></li>\n<li><p><strong>Final Validation Levenshtein Distance</strong></p>\n<ul>\n<li><strong><em>3.75</em></strong> <em>(not sure why this is different than LSD after epoch 12…)</em></li></ul></li>\n<li><p><strong>Public Leaderboard Score</strong></p>\n<ul>\n<li><strong><em>4.97</em></strong></li></ul></li>\n</ul>\n<p><em>SIDE NOTE: If you fine-tune on the training dataset for another 3-5 epochs with the max length being the correct value, I would estimate an improvement of 0.75-1.5 LSD on the LB… just a guestimate from previous experience though…</em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong><em>NOTE:</em></strong></p>\n<p>The discrepancy between public LB LSD and validation LSD is to be expected for 2 reasons.</p>\n<ol>\n<li>The validation LSD is calculated assuming a max length equivalent to half of the actual max length</li>\n<li>The strategy in place currently to handle the rotation for test images is not sufficient. I apply the h&gt;w then rotate strategy… but it fails to capture many of the cases for which rotation is required. </li>\n</ol>\n<p><br></p>\n<hr>\n<p><br></p>\n<p>I hope this helps!</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong><em>p.s.</em></strong> <em>I will be updating the notebook tomorrow or early this week to include an option to train the EfficientNetV2 model from the provided pre-trained ImageNet weights.</em></p>\n<p><strong><em>p.p.s.</em></strong> <em>I will also tack on the RDKit normalization procedure to the end of this notebook as per <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>'s notebook. Not sure if it will help… but why not.</em></p>",
      "rawMarkdown": "Hi there, \n\nI have created a notebook that implements the EfficientNetV2 model architecture (in Tensorflow) to function as the image encoder.\n\nNot only that. I implement it in such a way that **BOTH THE TRAINING AND THE MODEL INFERENCE ON THE TEST SET** can happen in **APPROXIMATELY 3 HOURS (OR LESS!)** \n\n---\n\n<br>\n\n***TL;DR  –––––––  TF+TPU+EfficientNetV2  –––>* 🔥🔥🔥🔥🔥🔥🔥🔥**\n\n<br>\n\n---\n\n*A caveat to the work I've completed. As this notebook implements both of these things, I don't have a comparative version with EfficientNetB2 to benchmark against. If I have time in the future I will do this just so we can all see the comparative performance differences.*\n\n*A second caveat is that I did the notebook development and initial submission within an interactive session. Therefore, there MAY be problems with a save-run-all scenario... I hope not. If I have TPU near the end of the week I will re-run this notebook using the run-all save option. If someone else does this and it succeeds with no alterations please let me know. Thanks* 😊\n\n<br>\n\n---\n\n<br>\n\n### <b>Links</b>\n* [**My End-to-End Pipeline Notebook**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/edit)\n* [**My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2**](https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2)\n* [**How To Load ImageNet Weights Into EfficientNetV2 And Infer** ***(for fine-tuning)***](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights)\n* [**The EfficientNetV2 Paper**](https://arxiv.org/pdf/2104.00298.pdf)\n* [**The Official Google Implementation of EfficientNetV2**](https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2)\n* [**The Show, Attend, and Tell Paper**](https://arxiv.org/pdf/1502.03044)\n\n<br>\n\n### <b>Model Details</b>\n\n* **Encoder**: \n  * ***EfficientNetV2-B2** (no pre-trained weights loaded)*\n* **Attention:**\n  * *[Bhadanau](https://arxiv.org/abs/1409.0473) w/ 192 Units*\n* **Decoder:** \n  * *LSTM w/ 512 Units and 0.05 Dropout*\n\n<br>\n\n### <b>Data Details</b>\n\n* **Data Size:**\n  * *Val -> 80,000 (random sample)*\n  * *Train -> All-80,000*\n  * *Test -> All*\n* **Batch Size:**\n  * *64 Per Replica (512 overall) ... this could have been higher in hindsight (128/1024)*\n* **Image Size:**\n  * *192x384x3*\n* **Image Prep:**\n  * *TRAIN: Crop To Molecule, Invert and Resize (no padding)*\n  * *TEST : Crop To Molecule, Invert and Resize (no padding) ... **also if h>w rotate 90 degrees***\n* **InChI Manipulation:**\n  * *Limit Max Length to Be Half Of Actual Max Length*\n* **InChI Tokenization:**\n  * *Entities Are Considered Tokens (i.e. `Br` is a token, `191` is a token, etc.)*\n\n<br>\n\n### <b>Pipeline and General Details</b>\n\n* **Execution & Development Environment**\n  * *Kaggle Notebook*\n* **Framework**\n  * *Tensorflow*\n* **Accelerator**\n  * *TPU v3-8*\n* **Data Format**\n  * *TFRecords*\n* **I/O Streaming Tools**\n  * *tf.data*\n* **Learning Rate Scheduler**\n  * *Ramp up to **2e-3**... step down to **1e-5***\n\n<br>\n\n### <b>Result Details</b>\n \n* **Timing**\n  * *~2-2.5 hours to complete training*\n    * *~10-12 minutes per epoch (includes validation)*\n    * *~53 seconds to infer on validation dataset afterwards for visualization's sake...*\n  * *14 minutes and 15 seconds for the entire inference and submission process*\n    * *5 minutes and 40 seconds to infer on test data*\n    * *8 minutes and 32 seconds to post-process test data (tokens to strings)*\n  * ***~2.75-3.25 hours to run the entire notebook including extras and visualizations***\n    * ***~2.5-2.75 hours of required computation... the rest is for show or display purposes***\n\n* **Validation Levenshtein Distance (up to half max length) By Epoch**\n  * *1  -> 15.79*\n  * *2  -> 9.51*\n  * *3  -> 6.99*\n  * *4  -> 6.74*\n  * *5  -> 4.91*\n  * *6  -> 4.88*\n  * *7  -> 4.02*\n  * *8  -> 4.01*\n  * *9  -> 3.64* \n  * *10 -> 3.77*\n  * *11 -> 3.42*\n  * *12 -> 3.32*\n* **Final Validation Levenshtein Distance**\n  * ***3.75*** *(not sure why this is different than LSD after epoch 12...)*\n* **Public Leaderboard Score**\n  * ***4.97***\n\n*SIDE NOTE: If you fine-tune on the training dataset for another 3-5 epochs with the max length being the correct value, I would estimate an improvement of 0.75-1.5 LSD on the LB... just a guestimate from previous experience though...*\n\n<br>\n\n---\n\n<br>\n\n***NOTE:***\n\nThe discrepancy between public LB LSD and validation LSD is to be expected for 2 reasons.\n\n1. The validation LSD is calculated assuming a max length equivalent to half of the actual max length\n2. The strategy in place currently to handle the rotation for test images is not sufficient. I apply the h>w then rotate strategy... but it fails to capture many of the cases for which rotation is required. \n\n<br>\n\n---\n\n<br>\n\nI hope this helps!\n\n<br>\n\n---\n\n<br>\n\n***p.s.*** *I will be updating the notebook tomorrow or early this week to include an option to train the EfficientNetV2 model from the provided pre-trained ImageNet weights.*\n\n***p.p.s.*** *I will also tack on the RDKit normalization procedure to the end of this notebook as per @nofreewill's notebook. Not sure if it will help... but why not.*",
      "votes": null
    },
    {
      "id": "1316322",
      "postDate": "05/20/2021 13:02:33",
      "content": "<p>I am currently trying a similar strategy of tuning on shorter sequences and then fine-tuning on full length. I see that your performance improved after choosing longer sequences but mine converges to a worse point. Is ther anything you tried to allow the model to converge better on the full sequences? Thanks!</p>",
      "rawMarkdown": "I am currently trying a similar strategy of tuning on shorter sequences and then fine-tuning on full length. I see that your performance improved after choosing longer sequences but mine converges to a worse point. Is ther anything you tried to allow the model to converge better on the full sequences? Thanks!",
      "votes": null
    },
    {
      "id": "1316412",
      "postDate": "05/20/2021 14:21:51",
      "content": "<p>You have to be very careful with the learning rate schedule. i.e. long ramp-ups and lower peak.<br>\nOther than that I have not done anything fancy. </p>\n<p>One note:</p>\n<p>I do not CHOOSE longer or shorter sequences. I simply truncate all of the sequences to a smaller fixed max length for the initial rounds of training and then remove that truncation for fine-tuning. I have not done it with the model above… but using an EfficientNetB5 we achieved an improvement in the order of 0.75-1.25 LSD by fine-tuning on the non-truncated sequences.</p>",
      "rawMarkdown": "You have to be very careful with the learning rate schedule. i.e. long ramp-ups and lower peak.\nOther than that I have not done anything fancy. \n\nOne note:\n\nI do not CHOOSE longer or shorter sequences. I simply truncate all of the sequences to a smaller fixed max length for the initial rounds of training and then remove that truncation for fine-tuning. I have not done it with the model above... but using an EfficientNetB5 we achieved an improvement in the order of 0.75-1.25 LSD by fine-tuning on the non-truncated sequences.",
      "votes": null
    },
    {
      "id": "1316429",
      "postDate": "05/20/2021 14:40:34",
      "content": "<p>I see! I did filter out the longer sequences. Perhaps my model biased and then lost many features after adding back in longer sequences. Thanks for all the help.</p>",
      "rawMarkdown": "I see! I did filter out the longer sequences. Perhaps my model biased and then lost many features after adding back in longer sequences. Thanks for all the help.",
      "votes": null
    },
    {
      "id": "1316468",
      "postDate": "05/20/2021 15:27:16",
      "content": "<p>Great notebook. I also tried training from scratch but I found out training from pretrained weight helped a lot. You are right that rotation if h &gt; w does not perform quite well, one example in the test data set is d83bc1c3b5ec. This image has h&gt;w , but in fact it should not be rotated based on the element orientation.   </p>",
      "rawMarkdown": "Great notebook. I also tried training from scratch but I found out training from pretrained weight helped a lot. You are right that rotation if h > w does not perform quite well, one example in the test data set is d83bc1c3b5ec. This image has h>w , but in fact it should not be rotated based on the element orientation.",
      "votes": null
    },
    {
      "id": "1316487",
      "postDate": "05/20/2021 15:33:59",
      "content": "<p>Really! Can you share a benchmark with starting from pretrained?</p>",
      "rawMarkdown": "Really! Can you share a benchmark with starting from pretrained?",
      "votes": null
    },
    {
      "id": "1317123",
      "postDate": "05/21/2021 07:03:20",
      "content": "<p>Great job! A question: you have a PREFIX_ORDERING which is not used anywhere - what is the purpose of it?</p>",
      "rawMarkdown": "Great job! A question: you have a PREFIX_ORDERING which is not used anywhere - what is the purpose of it?",
      "votes": null
    },
    {
      "id": "1317539",
      "postDate": "05/21/2021 13:31:43",
      "content": "<p>It’s a vestigial thing I created when I was doing my EDA. It just shows the order that the InChI sub layers go in.</p>\n<p>I don’t use it anywhere. Safe to delete!</p>",
      "rawMarkdown": "It’s a vestigial thing I created when I was doing my EDA. It just shows the order that the InChI sub layers go in.\n\nI don’t use it anywhere. Safe to delete!",
      "votes": null
    },
    {
      "id": "1318960",
      "postDate": "05/22/2021 17:54:58",
      "content": "<p>Hey! Great effort, just a question, which notebook did you use for generatin the tfrecords??<br>\nThanks</p>",
      "rawMarkdown": "Hey! Great effort, just a question, which notebook did you use for generatin the tfrecords??\nThanks",
      "votes": null
    },
    {
      "id": "1322907",
      "postDate": "05/25/2021 19:22:21",
      "content": "<p>See here -&gt; <a href=\"https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half</a></p>",
      "rawMarkdown": "See here -> https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half",
      "votes": null
    },
    {
      "id": "1322993",
      "postDate": "05/25/2021 21:19:15",
      "content": "<p>404 notebook is private or does not exist, but appreciate the answer, your other notebooks are following the same process to generate the tfrecords, thanks!</p>",
      "rawMarkdown": "404 notebook is private or does not exist, but appreciate the answer, your other notebooks are following the same process to generate the tfrecords, thanks!",
      "votes": null
    },
    {
      "id": "1323060",
      "postDate": "05/25/2021 23:46:38",
      "content": "<p>Updated it to be public! My bad!</p>",
      "rawMarkdown": "Updated it to be public! My bad!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1316322,
      "author_name": "andrewshao05",
      "author_url": "",
      "post_date": "05/20/2021 13:02:33",
      "content": "<p>I am currently trying a similar strategy of tuning on shorter sequences and then fine-tuning on full length. I see that your performance improved after choosing longer sequences but mine converges to a worse point. Is ther anything you tried to allow the model to converge better on the full sequences? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1316412,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/20/2021 14:21:51",
          "content": "<p>You have to be very careful with the learning rate schedule. i.e. long ramp-ups and lower peak.<br>\nOther than that I have not done anything fancy. </p>\n<p>One note:</p>\n<p>I do not CHOOSE longer or shorter sequences. I simply truncate all of the sequences to a smaller fixed max length for the initial rounds of training and then remove that truncation for fine-tuning. I have not done it with the model above… but using an EfficientNetB5 we achieved an improvement in the order of 0.75-1.25 LSD by fine-tuning on the non-truncated sequences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316429,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/20/2021 14:40:34",
          "content": "<p>I see! I did filter out the longer sequences. Perhaps my model biased and then lost many features after adding back in longer sequences. Thanks for all the help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316468,
      "author_name": "joblessphysicist",
      "author_url": "",
      "post_date": "05/20/2021 15:27:16",
      "content": "<p>Great notebook. I also tried training from scratch but I found out training from pretrained weight helped a lot. You are right that rotation if h &gt; w does not perform quite well, one example in the test data set is d83bc1c3b5ec. This image has h&gt;w , but in fact it should not be rotated based on the element orientation.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 1316487,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/20/2021 15:33:59",
          "content": "<p>Really! Can you share a benchmark with starting from pretrained?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1317123,
      "author_name": "jhasanov",
      "author_url": "",
      "post_date": "05/21/2021 07:03:20",
      "content": "<p>Great job! A question: you have a PREFIX_ORDERING which is not used anywhere - what is the purpose of it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1317539,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/21/2021 13:31:43",
          "content": "<p>It’s a vestigial thing I created when I was doing my EDA. It just shows the order that the InChI sub layers go in.</p>\n<p>I don’t use it anywhere. Safe to delete!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1318960,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "05/22/2021 17:54:58",
      "content": "<p>Hey! Great effort, just a question, which notebook did you use for generatin the tfrecords??<br>\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1322907,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/25/2021 19:22:21",
          "content": "<p>See here -&gt; <a href=\"https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1322993,
          "author_name": "jesucristo",
          "author_url": "",
          "post_date": "05/25/2021 21:19:15",
          "content": "<p>404 notebook is private or does not exist, but appreciate the answer, your other notebooks are following the same process to generate the tfrecords, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323060,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/25/2021 23:46:38",
          "content": "<p>Updated it to be public! My bad!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1310687": "Hi there, \n\nI have created a notebook that implements the EfficientNetV2 model architecture (in Tensorflow) to function as the image encoder.\n\nNot only that. I implement it in such a way that **BOTH THE TRAINING AND THE MODEL INFERENCE ON THE TEST SET** can happen in **APPROXIMATELY 3 HOURS (OR LESS!)** \n\n---\n\n<br>\n\n***TL;DR  –––––––  TF+TPU+EfficientNetV2  –––>* 🔥🔥🔥🔥🔥🔥🔥🔥**\n\n<br>\n\n---\n\n*A caveat to the work I've completed. As this notebook implements both of these things, I don't have a comparative version with EfficientNetB2 to benchmark against. If I have time in the future I will do this just so we can all see the comparative performance differences.*\n\n*A second caveat is that I did the notebook development and initial submission within an interactive session. Therefore, there MAY be problems with a save-run-all scenario... I hope not. If I have TPU near the end of the week I will re-run this notebook using the run-all save option. If someone else does this and it succeeds with no alterations please let me know. Thanks* 😊\n\n<br>\n\n---\n\n<br>\n\n### <b>Links</b>\n* [**My End-to-End Pipeline Notebook**](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs/edit)\n* [**My Dataset Containing A Modified Version of the Official Github Repo For EfficientNetV2**](https://www.kaggle.com/dschettler8845/automl-efficientdet-efficientnetv2)\n* [**How To Load ImageNet Weights Into EfficientNetV2 And Infer** ***(for fine-tuning)***](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights)\n* [**The EfficientNetV2 Paper**](https://arxiv.org/pdf/2104.00298.pdf)\n* [**The Official Google Implementation of EfficientNetV2**](https://github.com/google/automl/tree/4d2c9286d3ff7ba5b3a231881746274ddf52c031/efficientnetv2)\n* [**The Show, Attend, and Tell Paper**](https://arxiv.org/pdf/1502.03044)\n\n<br>\n\n### <b>Model Details</b>\n\n* **Encoder**: \n  * ***EfficientNetV2-B2** (no pre-trained weights loaded)*\n* **Attention:**\n  * *[Bhadanau](https://arxiv.org/abs/1409.0473) w/ 192 Units*\n* **Decoder:** \n  * *LSTM w/ 512 Units and 0.05 Dropout*\n\n<br>\n\n### <b>Data Details</b>\n\n* **Data Size:**\n  * *Val -> 80,000 (random sample)*\n  * *Train -> All-80,000*\n  * *Test -> All*\n* **Batch Size:**\n  * *64 Per Replica (512 overall) ... this could have been higher in hindsight (128/1024)*\n* **Image Size:**\n  * *192x384x3*\n* **Image Prep:**\n  * *TRAIN: Crop To Molecule, Invert and Resize (no padding)*\n  * *TEST : Crop To Molecule, Invert and Resize (no padding) ... **also if h>w rotate 90 degrees***\n* **InChI Manipulation:**\n  * *Limit Max Length to Be Half Of Actual Max Length*\n* **InChI Tokenization:**\n  * *Entities Are Considered Tokens (i.e. `Br` is a token, `191` is a token, etc.)*\n\n<br>\n\n### <b>Pipeline and General Details</b>\n\n* **Execution & Development Environment**\n  * *Kaggle Notebook*\n* **Framework**\n  * *Tensorflow*\n* **Accelerator**\n  * *TPU v3-8*\n* **Data Format**\n  * *TFRecords*\n* **I/O Streaming Tools**\n  * *tf.data*\n* **Learning Rate Scheduler**\n  * *Ramp up to **2e-3**... step down to **1e-5***\n\n<br>\n\n### <b>Result Details</b>\n \n* **Timing**\n  * *~2-2.5 hours to complete training*\n    * *~10-12 minutes per epoch (includes validation)*\n    * *~53 seconds to infer on validation dataset afterwards for visualization's sake...*\n  * *14 minutes and 15 seconds for the entire inference and submission process*\n    * *5 minutes and 40 seconds to infer on test data*\n    * *8 minutes and 32 seconds to post-process test data (tokens to strings)*\n  * ***~2.75-3.25 hours to run the entire notebook including extras and visualizations***\n    * ***~2.5-2.75 hours of required computation... the rest is for show or display purposes***\n\n* **Validation Levenshtein Distance (up to half max length) By Epoch**\n  * *1  -> 15.79*\n  * *2  -> 9.51*\n  * *3  -> 6.99*\n  * *4  -> 6.74*\n  * *5  -> 4.91*\n  * *6  -> 4.88*\n  * *7  -> 4.02*\n  * *8  -> 4.01*\n  * *9  -> 3.64* \n  * *10 -> 3.77*\n  * *11 -> 3.42*\n  * *12 -> 3.32*\n* **Final Validation Levenshtein Distance**\n  * ***3.75*** *(not sure why this is different than LSD after epoch 12...)*\n* **Public Leaderboard Score**\n  * ***4.97***\n\n*SIDE NOTE: If you fine-tune on the training dataset for another 3-5 epochs with the max length being the correct value, I would estimate an improvement of 0.75-1.5 LSD on the LB... just a guestimate from previous experience though...*\n\n<br>\n\n---\n\n<br>\n\n***NOTE:***\n\nThe discrepancy between public LB LSD and validation LSD is to be expected for 2 reasons.\n\n1. The validation LSD is calculated assuming a max length equivalent to half of the actual max length\n2. The strategy in place currently to handle the rotation for test images is not sufficient. I apply the h>w then rotate strategy... but it fails to capture many of the cases for which rotation is required. \n\n<br>\n\n---\n\n<br>\n\nI hope this helps!\n\n<br>\n\n---\n\n<br>\n\n***p.s.*** *I will be updating the notebook tomorrow or early this week to include an option to train the EfficientNetV2 model from the provided pre-trained ImageNet weights.*\n\n***p.p.s.*** *I will also tack on the RDKit normalization procedure to the end of this notebook as per @nofreewill's notebook. Not sure if it will help... but why not.*",
    "1316322": "I am currently trying a similar strategy of tuning on shorter sequences and then fine-tuning on full length. I see that your performance improved after choosing longer sequences but mine converges to a worse point. Is ther anything you tried to allow the model to converge better on the full sequences? Thanks!",
    "1316412": "You have to be very careful with the learning rate schedule. i.e. long ramp-ups and lower peak.\nOther than that I have not done anything fancy. \n\nOne note:\n\nI do not CHOOSE longer or shorter sequences. I simply truncate all of the sequences to a smaller fixed max length for the initial rounds of training and then remove that truncation for fine-tuning. I have not done it with the model above... but using an EfficientNetB5 we achieved an improvement in the order of 0.75-1.25 LSD by fine-tuning on the non-truncated sequences.",
    "1316429": "I see! I did filter out the longer sequences. Perhaps my model biased and then lost many features after adding back in longer sequences. Thanks for all the help.",
    "1316468": "Great notebook. I also tried training from scratch but I found out training from pretrained weight helped a lot. You are right that rotation if h > w does not perform quite well, one example in the test data set is d83bc1c3b5ec. This image has h>w , but in fact it should not be rotated based on the element orientation.",
    "1316487": "Really! Can you share a benchmark with starting from pretrained?",
    "1317123": "Great job! A question: you have a PREFIX_ORDERING which is not used anywhere - what is the purpose of it?",
    "1317539": "It’s a vestigial thing I created when I was doing my EDA. It just shows the order that the InChI sub layers go in.\n\nI don’t use it anywhere. Safe to delete!",
    "1318960": "Hey! Great effort, just a question, which notebook did you use for generatin the tfrecords??\nThanks",
    "1322907": "See here -> https://www.kaggle.com/dschettler8845/bms-tfrecord-creation-192x384-train-half",
    "1322993": "404 notebook is private or does not exist, but appreciate the answer, your other notebooks are following the same process to generate the tfrecords, thanks!",
    "1323060": "Updated it to be public! My bad!"
  },
  "source": "meta"
}