{
  "id": 239236,
  "title": "Pretrained Weights and code for Efficientnet V2 Models",
  "url": "/competitions/bms-molecular-translation/discussion/239236",
  "author_name": "Saurabh Shahane",
  "post_date": "2021-05-15T11:51:30.205000",
  "votes": 10,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Full code and pretrained weights for efficientnet v2 can be found here - <a href=\"https://github.com/google/automl/tree/master/efficientnetv2\" target=\"_blank\">https://github.com/google/automl/tree/master/efficientnetv2</a></p>\n<p><img src=\"https://github.com/google/automl/raw/master/efficientnetv2/g3doc/train_params.png\" alt=\"Efficientnet Weights\"></p>",
  "messages": [
    {
      "id": 1308697,
      "postDate": "2021-05-15T11:51:30.207Z",
      "content": "<p>Full code and pretrained weights for efficientnet v2 can be found here - <a href=\"https://github.com/google/automl/tree/master/efficientnetv2\" target=\"_blank\">https://github.com/google/automl/tree/master/efficientnetv2</a></p>\n<p><img src=\"https://github.com/google/automl/raw/master/efficientnetv2/g3doc/train_params.png\" alt=\"Efficientnet Weights\"></p>",
      "rawMarkdown": "Full code and pretrained weights for efficientnet v2 can be found here - https://github.com/google/automl/tree/master/efficientnetv2\n\n\n\n![Efficientnet Weights](https://github.com/google/automl/raw/master/efficientnetv2/g3doc/train_params.png)",
      "votes": 10
    },
    {
      "id": 1310691,
      "postDate": "2021-05-16T21:09:11.733Z",
      "content": "<p>End-to-end pipeline including training, inference on the test set, and submission using an EfficientNetV2 is now complete!! The whole thing takes less than 3 hours.</p>\n<ul>\n<li><strong>Resulting score is 4.97 Levenshtein Distance on the public LB.</strong></li>\n</ul>\n<hr>\n<p>Here is the <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">notebook</a></p>\n<p>Here is my <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/239595\" target=\"_blank\">discussion post</a> which links to all the relevant stuff.</p>\n<hr>\n<p>Side Note: Here is a <a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\">notebook</a> I created showing how to load the pre-trained weights. I do inference on a random photo to show that it works.</p>\n<hr>\n<p>Hope this helps!</p>",
      "rawMarkdown": "End-to-end pipeline including training, inference on the test set, and submission using an EfficientNetV2 is now complete!! The whole thing takes less than 3 hours.\n* **Resulting score is 4.97 Levenshtein Distance on the public LB.**\n\n---\n\nHere is the [notebook](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)\n\nHere is my [discussion post](https://www.kaggle.com/c/bms-molecular-translation/discussion/239595) which links to all the relevant stuff.\n\n---\n\nSide Note: Here is a [notebook](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights) I created showing how to load the pre-trained weights. I do inference on a random photo to show that it works.\n\n---\n\nHope this helps!",
      "votes": 3,
      "replies": [
        {
          "id": 1320327,
          "postDate": "2021-05-24T02:25:36.880Z",
          "content": "<p>could you please give a notebook example for transfer learning ? I tried to build one using Keras model but not succeeded.  thanks!</p>",
          "rawMarkdown": "could you please give a notebook example for transfer learning ? I tried to build one using Keras model but not succeeded.  thanks!"
        }
      ]
    },
    {
      "id": 1309407,
      "postDate": "2021-05-15T23:04:18.483Z",
      "content": "<p>Will implement and share a notebook leveraging keras and tpu this weekend. <strong>[UPDATE: DONE]</strong></p>",
      "rawMarkdown": "Will implement and share a notebook leveraging keras and tpu this weekend. **[UPDATE: DONE]**",
      "votes": 3,
      "replies": [
        {
          "id": 1309715,
          "postDate": "2021-05-16T08:23:59.180Z",
          "content": "<p>Looking forward to it <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> !!</p>",
          "rawMarkdown": "Looking forward to it @dschettler8845 !!",
          "votes": 1
        },
        {
          "id": 1310695,
          "postDate": "2021-05-16T21:10:44.917Z",
          "content": "<p>See my reply above. I have created a notebook, discussion post, and dataset. Hope this helps!</p>",
          "rawMarkdown": "See my reply above. I have created a notebook, discussion post, and dataset. Hope this helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1309746,
      "postDate": "2021-05-16T08:55:10.837Z",
      "content": "<p>It seems the default weight is the pretrained Imagenet1k </p>\n<p>However the bigger Imagenet (21k + finetune on 1k) give better results on the paper.   I'm experimenting with this last one to see (a bit trickier to load the weight in the context of distributed training)</p>",
      "rawMarkdown": "It seems the default weight is the pretrained Imagenet1k \n\nHowever the bigger Imagenet (21k + finetune on 1k) give better results on the paper.   I'm experimenting with this last one to see (a bit trickier to load the weight in the context of distributed training)",
      "votes": 1
    },
    {
      "id": 1312980,
      "postDate": "2021-05-18T11:02:30.703Z",
      "content": "<p>Based on my first expriments,  EffnetV1  with noisy student weight give better results in this dataset than EFFnetV2.</p>\n<p>I think EFFnetV2 would benefit from NS training given how it improved significantly EffnetV1. </p>\n<p>May be the comparaison in the paper is with EffnetV1 without NS training. </p>",
      "rawMarkdown": "Based on my first expriments,  EffnetV1  with noisy student weight give better results in this dataset than EFFnetV2.\n\nI think EFFnetV2 would benefit from NS training given how it improved significantly EffnetV1. \n\nMay be the comparaison in the paper is with EffnetV1 without NS training. "
    },
    {
      "id": 1310637,
      "postDate": "2021-05-16T19:23:09.097Z",
      "content": "<p>Happy to hear the models and weights are finally made available.<br>\nI did take a look at it, however there does not seem to be an option to load the pretrained weights when building the model.</p>\n<pre><code># build keras model\nmodel = effnetv2_model.EffNetV2Model('efficientnetv2-s')\n# how to load the pretrained weights?\n...\n</code></pre>\n<p>When unzipping the checkpoints there are 4 files:</p>\n<pre><code>checkpoint\nmodel.data-00000-of-00001\nmodel.index\nmodel.meta\n</code></pre>\n<p>Conventionally the weights are given as a single h5 file, how to load the given weights?</p>",
      "rawMarkdown": "Happy to hear the models and weights are finally made available.\nI did take a look at it, however there does not seem to be an option to load the pretrained weights when building the model.\n\n```\n# build keras model\nmodel = effnetv2_model.EffNetV2Model('efficientnetv2-s')\n# how to load the pretrained weights?\n...\n```\n\nWhen unzipping the checkpoints there are 4 files:\n\n```\ncheckpoint\nmodel.data-00000-of-00001\nmodel.index\nmodel.meta\n```\n\nConventionally the weights are given as a single h5 file, how to load the given weights?",
      "replies": [
        {
          "id": 1310640,
          "postDate": "2021-05-16T19:26:26.320Z",
          "content": "<p>They provide the weights as ckpts. I have not implemented this part yet (as I’m focused on training from scratch… the notebook and post will be shared within 2 hours). </p>\n<p>I assume you could build the model and load the weights by name maybe? I’m not sure though…</p>\n<p>That being said I will make it a priority to demonstrate that functionality… probably tomorrow.</p>",
          "rawMarkdown": "They provide the weights as ckpts. I have not implemented this part yet (as I’m focused on training from scratch... the notebook and post will be shared within 2 hours). \n\nI assume you could build the model and load the weights by name maybe? I’m not sure though...\n\nThat being said I will make it a priority to demonstrate that functionality... probably tomorrow.",
          "votes": 1
        },
        {
          "id": 1310670,
          "postDate": "2021-05-16T20:15:14.137Z",
          "content": "<p>I did get a small step further, the <code>EffNetV2Model</code> extends <code>tf.keras.Model</code> and does therefore support the <code>load_weights</code> functions. However, an error is thrown when the checkpoint is loaded this way.</p>\n<pre><code>m.load_weights('path_to_weights/efficientnetv2-s-21k-ft1k/model')\n</code></pre>\n<p>Output:</p>\n<pre><code>AssertionError: Some objects had attributes which were not restored:\n\n[A huge amount of layers which can't be loaded...]\n</code></pre>\n<p>The tutorial I followed for loading the checkpoint can be found <a href=\"https://www.tensorflow.org/tutorials/keras/save_and_load#checkpoint_callback_usage\" target=\"_blank\">here</a></p>\n<p>Happy to hear when you get any progress.</p>",
          "rawMarkdown": "I did get a small step further, the `EffNetV2Model` extends `tf.keras.Model` and does therefore support the `load_weights` functions. However, an error is thrown when the checkpoint is loaded this way.\n\n```\nm.load_weights('path_to_weights/efficientnetv2-s-21k-ft1k/model')\n```\n\nOutput:\n\n```\nAssertionError: Some objects had attributes which were not restored:\n\n[A huge amount of layers which can't be loaded...]\n```\n\nThe tutorial I followed for loading the checkpoint can be found [here](https://www.tensorflow.org/tutorials/keras/save_and_load#checkpoint_callback_usage)\n\nHappy to hear when you get any progress.",
          "votes": 1
        },
        {
          "id": 1310696,
          "postDate": "2021-05-16T21:11:07.803Z",
          "content": "<p>The weight in 4 files is the default TF format (not keras). Thus you need to use the default TF checkpoint loader from directory: <code>tf.train.Checkpoint</code>  </p>\n<pre><code>model = effnetv2_model.EffNetV2Model('efficientnetv2-s')\nCKPT_dir ='directory_of_dowloaded_weight'\ncheckpoint = tf.train.Checkpoint(model)\ncheckpoint.restore(tf.train.latest_checkpoint(CKPT_dir ))\n</code></pre>",
          "rawMarkdown": "The weight in 4 files is the default TF format (not keras). Thus you need to use the default TF checkpoint loader from directory: `tf.train.Checkpoint`  \n\n\n```\nmodel = effnetv2_model.EffNetV2Model('efficientnetv2-s')\nCKPT_dir ='directory_of_dowloaded_weight'\ncheckpoint = tf.train.Checkpoint(model)\ncheckpoint.restore(tf.train.latest_checkpoint(CKPT_dir ))\n```\n\n\n\n",
          "votes": 1
        },
        {
          "id": 1310729,
          "postDate": "2021-05-16T22:18:27.790Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1310744,
          "postDate": "2021-05-16T22:53:01.850Z",
          "content": "<p>Did anyone get this to work. I'm still just getting: AssertionError: Some objects had attributes which were not restored, when checking the loading status, and every weight is listed there(I.e No weight is loaded, but they all exist).</p>",
          "rawMarkdown": "Did anyone get this to work. I'm still just getting: AssertionError: Some objects had attributes which were not restored, when checking the loading status, and every weight is listed there(I.e No weight is loaded, but they all exist)."
        },
        {
          "id": 1310790,
          "postDate": "2021-05-17T00:57:42.650Z",
          "content": "<p>I'm making a notebook now… but this is from the infer.py file within the repo… seems straightforward…</p>\n<pre><code>def build_tf2_model():\n  \"\"\"Build the tf2 model.\"\"\"\n  # Use 'mixed_float16' if running on GPUs.\n  policy = tf.keras.mixed_precision.Policy('mixed_float16')\n  tf.keras.mixed_precision.set_global_policy(policy)\n  tf.config.run_functions_eagerly(FLAGS.debug)\n  # Create and run the model.\n  model = create_model(FLAGS.model_name, FLAGS.dataset_cfg, FLAGS.hparam_str)\n  # Use call (not build) to match the namescope: tensorflow issues/29576\n  model(tf.ones([1, 224, 224, 3]), False)\n  if FLAGS.model_dir:\n    ckpt = FLAGS.model_dir\n    if tf.io.gfile.isdir(ckpt):\n      ckpt = tf.train.latest_checkpoint(FLAGS.model_dir)\n    model.load_weights(ckpt)\n  model.summary()\n</code></pre>",
          "rawMarkdown": "I'm making a notebook now... but this is from the infer.py file within the repo... seems straightforward...\n\n```python\ndef build_tf2_model():\n  \"\"\"Build the tf2 model.\"\"\"\n  # Use 'mixed_float16' if running on GPUs.\n  policy = tf.keras.mixed_precision.Policy('mixed_float16')\n  tf.keras.mixed_precision.set_global_policy(policy)\n  tf.config.run_functions_eagerly(FLAGS.debug)\n  # Create and run the model.\n  model = create_model(FLAGS.model_name, FLAGS.dataset_cfg, FLAGS.hparam_str)\n  # Use call (not build) to match the namescope: tensorflow issues/29576\n  model(tf.ones([1, 224, 224, 3]), False)\n  if FLAGS.model_dir:\n    ckpt = FLAGS.model_dir\n    if tf.io.gfile.isdir(ckpt):\n      ckpt = tf.train.latest_checkpoint(FLAGS.model_dir)\n    model.load_weights(ckpt)\n  model.summary()\n```"
        },
        {
          "id": 1310796,
          "postDate": "2021-05-17T01:08:04.013Z",
          "content": "<p>Here… I made it into a notebook. I load the weights and infer on a random image of a cat.</p>\n<p><strong><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\">https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights</a></strong></p>",
          "rawMarkdown": "Here... I made it into a notebook. I load the weights and infer on a random image of a cat.\n\n**https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights**",
          "votes": 1
        },
        {
          "id": 1310799,
          "postDate": "2021-05-17T01:14:24.740Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1310968,
          "postDate": "2021-05-17T04:55:34.250Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrewshao05\" target=\"_blank\">@andrewshao05</a> The snippet I give, works for me.  I am training right now with the pretrained 21k + finetune 1k weight. </p>\n<p>It may be need some cautious with your optimizer. </p>",
          "rawMarkdown": "@andrewshao05 The snippet I give, works for me.  I am training right now with the pretrained 21k + finetune 1k weight. \n\nIt may be need some cautious with your optimizer. "
        },
        {
          "id": 1310983,
          "postDate": "2021-05-17T05:02:51.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>   Does this approach work for finetuning (not inference) ? </p>\n<p>In general, I think TF checkpoint management is unnecessarily complicated and misleading with too many formats,  while in Pytorch everything is simple given you can save and load straightforwardly all states dict (last epoch, lr_schedule, model and optimizer weights) with the same single file. </p>\n<p>That's why I prefer to use plain numpy arrays and manage my own checkpoint during training and use neither TF checkpoint nor keras default format. </p>",
          "rawMarkdown": "@dschettler8845   Does this approach work for finetuning (not inference) ? \n\nIn general, I think TF checkpoint management is unnecessarily complicated and misleading with too many formats,  while in Pytorch everything is simple given you can save and load straightforwardly all states dict (last epoch, lr_schedule, model and optimizer weights) with the same single file. \n\n\nThat's why I prefer to use plain numpy arrays and manage my own checkpoint during training and use neither TF checkpoint nor keras default format. "
        },
        {
          "id": 1329232,
          "postDate": "2021-05-31T01:40:35.640Z",
          "content": "<p>InvalidArgumentError: Unsuccessful TensorSliceReader constructor: Failed to get matching files on /content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model: Unimplemented: File system scheme '[local]' not implemented (file: '/content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model')</p>\n<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> after put your codes why I get error like this? Can you please suggests me what was the problem?</p>",
          "rawMarkdown": "InvalidArgumentError: Unsuccessful TensorSliceReader constructor: Failed to get matching files on /content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model: Unimplemented: File system scheme '[local]' not implemented (file: '/content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model')\n\n@serigne after put your codes why I get error like this? Can you please suggests me what was the problem?"
        },
        {
          "id": 1329919,
          "postDate": "2021-05-31T13:21:30.873Z",
          "content": "<p>If you are using TPU's this error can occur when using dataset paths instead of Google Cloud paths. Make your dataset public and use the Google Cloud path to the weights.</p>\n<pre><code>GOOGLE_CLOUD_PATH= KaggleDatasets().get_gcs_path('your-dataset')\n\n# load model\nmodel= effnetv2_model.EffNetV2Model(model_name='efficientnetv2-b3')\n# dummy call to initialize model\nmodel(tf.ones((1,224,224,3)), training=False)\n\n# weight path on Google Cloud\nWEIGHT_PATH = f'{GOOGLE_CLOUD_PATH}/efficientnetv2-b3'\n# Get the latest checkpoint from path\nckpt = tf.train.latest_checkpoint(WEIGHT_PATH)\n\n# Load the weights\nmodel.load_weights(ckpt)\n</code></pre>",
          "rawMarkdown": "If you are using TPU's this error can occur when using dataset paths instead of Google Cloud paths. Make your dataset public and use the Google Cloud path to the weights.\n\n```\nGOOGLE_CLOUD_PATH= KaggleDatasets().get_gcs_path('your-dataset')\n\n# load model\nmodel= effnetv2_model.EffNetV2Model(model_name='efficientnetv2-b3')\n# dummy call to initialize model\nmodel(tf.ones((1,224,224,3)), training=False)\n\n# weight path on Google Cloud\nWEIGHT_PATH = f'{GOOGLE_CLOUD_PATH}/efficientnetv2-b3'\n# Get the latest checkpoint from path\nckpt = tf.train.latest_checkpoint(WEIGHT_PATH)\n\n# Load the weights\nmodel.load_weights(ckpt)\n```",
          "votes": 2
        },
        {
          "id": 1330778,
          "postDate": "2021-06-01T05:07:32.297Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>.  Any suggestion for adding \"efficientnet-v2-s\" from \"efficientnetv2-s-21k\" I actually wrote \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s-21k\")\" but got error also I tried \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s\")\" but still got error.</p>",
          "rawMarkdown": "Thanks a lot @markwijkhuizen.  Any suggestion for adding \"efficientnet-v2-s\" from \"efficientnetv2-s-21k\" I actually wrote \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s-21k\")\" but got error also I tried \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s\")\" but still got error."
        },
        {
          "id": 1331351,
          "postDate": "2021-06-01T12:16:01.303Z",
          "content": "<p>For me using the following method works:</p>\n<p><code>model = effnetv2_model.EffNetV2Model(model_name='efficientnetv2-s')</code></p>\n<p>What kind of error do you get when loading the model?</p>",
          "rawMarkdown": "For me using the following method works:\n\n`model = effnetv2_model.EffNetV2Model(model_name='efficientnetv2-s')`\n\nWhat kind of error do you get when loading the model?"
        },
        {
          "id": 1331415,
          "postDate": "2021-06-01T13:06:10.117Z",
          "content": "<p>I found <strong>AssertionError</strong>:<br>\n&lt;AutoCastVariable 'efficientnetv2-s/stem/conv2d/kernel:0' shape=(3, 3, 3, 24) dtype=float32 dtype_to_cast_to=float32, numpy=<br>\narray([[[[ 1.31866440e-01,  2.92598866e-02,  5.63979149e-02,<br>\n           6.79737628e-02, -1.03196532e-01, -1.61898583e-02,<br>\n          -9.67781842e-02,  9.58853662e-02, -1.21669516e-01,<br>\n          -3.78544554e-02, -4.95341495e-02, -1.13189340e-01,<br>\n          -3.57197300e-02,  1.19056180e-01, -3.08847930e-02,<br>\n          -5.60378395e-02,  2.59503573e-02, -4.53319354e-03,<br>\n           3.76539188e-03, -2.62582731e-02, -4.90069948e-03,<br>\n          -2.31419038e-02,  1.32807806e-01,  1.46531919e-02],<br>\n         [ 2.10016251e-01, -5.64353578e-02,  2.44152751e-02,<br>\n           7.11440295e-02, -7.25007206e-02, -6.13373555e-02,<br>\n          -1.60715938e-01,  9.48695615e-02, -1.58229873e-01,<br>\n          -1.13155447e-01, -5.96799143e-02,  9.18575004e-02,<br>\n           2.68904697e-02, -2.05748565e-02, -1.37787834e-01,<br>\n          -4.67350520e-02, -7.78039470e-02,  7.96402842e-02,<br>\n           5.23473695e-02,  1.15151033e-01,  5.91931865e-03,<br>\n          -6.38899282e-02, -2.02910174e-02,  4.16407287e-02],<br>\n     ………………………………………………………………………………………….<br>\n     ………………………………………………………………………………………….</p>",
          "rawMarkdown": "I found **AssertionError**:\n<AutoCastVariable 'efficientnetv2-s/stem/conv2d/kernel:0' shape=(3, 3, 3, 24) dtype=float32 dtype_to_cast_to=float32, numpy=\narray([[[[ 1.31866440e-01,  2.92598866e-02,  5.63979149e-02,\n           6.79737628e-02, -1.03196532e-01, -1.61898583e-02,\n          -9.67781842e-02,  9.58853662e-02, -1.21669516e-01,\n          -3.78544554e-02, -4.95341495e-02, -1.13189340e-01,\n          -3.57197300e-02,  1.19056180e-01, -3.08847930e-02,\n          -5.60378395e-02,  2.59503573e-02, -4.53319354e-03,\n           3.76539188e-03, -2.62582731e-02, -4.90069948e-03,\n          -2.31419038e-02,  1.32807806e-01,  1.46531919e-02],\n         [ 2.10016251e-01, -5.64353578e-02,  2.44152751e-02,\n           7.11440295e-02, -7.25007206e-02, -6.13373555e-02,\n          -1.60715938e-01,  9.48695615e-02, -1.58229873e-01,\n          -1.13155447e-01, -5.96799143e-02,  9.18575004e-02,\n           2.68904697e-02, -2.05748565e-02, -1.37787834e-01,\n          -4.67350520e-02, -7.78039470e-02,  7.96402842e-02,\n           5.23473695e-02,  1.15151033e-01,  5.91931865e-03,\n          -6.38899282e-02, -2.02910174e-02,  4.16407287e-02],\n     .......................................................................................................\n     .......................................................................................................\n"
        },
        {
          "id": 1331419,
          "postDate": "2021-06-01T13:08:21.007Z",
          "content": "<p>I used for dataset directory from kaggle : </p>\n<blockquote>\n  <p>gs://kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-s-21k</p>\n</blockquote>\n<p>And when I try to load weights I found this error</p>",
          "rawMarkdown": "I used for dataset directory from kaggle : \n>gs://kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-s-21k\n\nAnd when I try to load weights I found this error"
        },
        {
          "id": 1331439,
          "postDate": "2021-06-01T13:19:57.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> - Looks like you are using the <strong><code>name</code></strong> argument instead of the <strong><code>model_name</code></strong>  argument. One is the string name associated with the keras model name (<strong><code>name</code></strong>) and the other informs on which backbone to load (<strong><code>model_name</code></strong>).</p>\n<p>Try changing that argument and seeing what happens?</p>",
          "rawMarkdown": "@aifahim - Looks like you are using the **`name`** argument instead of the **`model_name`**  argument. One is the string name associated with the keras model name (**`name`**) and the other informs on which backbone to load (**`model_name`**).\n\nTry changing that argument and seeing what happens?"
        },
        {
          "id": 1331450,
          "postDate": "2021-06-01T13:28:11.927Z",
          "content": "<blockquote>\n  <p>ev2_s = effnetv2_model.EffNetV2Model(model_name=\"efficientnetv2-s\", name=\"efficientnetv2-s\") </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> I actually used your code and change Ii with \"efficientnetv2-s\" from \"efficientnetv2-b0\"</p>",
          "rawMarkdown": ">ev2_s = effnetv2_model.EffNetV2Model(model_name=\"efficientnetv2-s\", name=\"efficientnetv2-s\") \n\n@dschettler8845 I actually used your code and change Ii with \"efficientnetv2-s\" from \"efficientnetv2-b0\""
        },
        {
          "id": 1332304,
          "postDate": "2021-06-02T03:31:51.547Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> - this method works for fine tuning. I’ll be open sourcing a notebook showing the same in the COVID competition shortly and will tag you when it’s complete.</p>",
          "rawMarkdown": "@serigne - this method works for fine tuning. I’ll be open sourcing a notebook showing the same in the COVID competition shortly and will tag you when it’s complete."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1310691,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-16T21:09:11.733000",
      "content": "<p>End-to-end pipeline including training, inference on the test set, and submission using an EfficientNetV2 is now complete!! The whole thing takes less than 3 hours.</p>\n<ul>\n<li><strong>Resulting score is 4.97 Levenshtein Distance on the public LB.</strong></li>\n</ul>\n<hr>\n<p>Here is the <a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">notebook</a></p>\n<p>Here is my <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/239595\" target=\"_blank\">discussion post</a> which links to all the relevant stuff.</p>\n<hr>\n<p>Side Note: Here is a <a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\">notebook</a> I created showing how to load the pre-trained weights. I do inference on a random photo to show that it works.</p>\n<hr>\n<p>Hope this helps!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1320327,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-05-24T02:25:36.880000",
          "content": "<p>could you please give a notebook example for transfer learning ? I tried to build one using Keras model but not succeeded.  thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1309407,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-15T23:04:18.483000",
      "content": "<p>Will implement and share a notebook leveraging keras and tpu this weekend. <strong>[UPDATE: DONE]</strong></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1309715,
          "author_name": "Saurabh Shahane",
          "author_url": "",
          "post_date": "2021-05-16T08:23:59.180000",
          "content": "<p>Looking forward to it <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> !!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1310695,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-16T21:10:44.917000",
          "content": "<p>See my reply above. I have created a notebook, discussion post, and dataset. Hope this helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1309746,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-05-16T08:55:10.837000",
      "content": "<p>It seems the default weight is the pretrained Imagenet1k </p>\n<p>However the bigger Imagenet (21k + finetune on 1k) give better results on the paper.   I'm experimenting with this last one to see (a bit trickier to load the weight in the context of distributed training)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1312980,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-05-18T11:02:30.703000",
      "content": "<p>Based on my first expriments,  EffnetV1  with noisy student weight give better results in this dataset than EFFnetV2.</p>\n<p>I think EFFnetV2 would benefit from NS training given how it improved significantly EffnetV1. </p>\n<p>May be the comparaison in the paper is with EffnetV1 without NS training. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1310637,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2021-05-16T19:23:09.097000",
      "content": "<p>Happy to hear the models and weights are finally made available.<br>\nI did take a look at it, however there does not seem to be an option to load the pretrained weights when building the model.</p>\n<pre><code># build keras model\nmodel = effnetv2_model.EffNetV2Model('efficientnetv2-s')\n# how to load the pretrained weights?\n...\n</code></pre>\n<p>When unzipping the checkpoints there are 4 files:</p>\n<pre><code>checkpoint\nmodel.data-00000-of-00001\nmodel.index\nmodel.meta\n</code></pre>\n<p>Conventionally the weights are given as a single h5 file, how to load the given weights?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1310640,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-16T19:26:26.320000",
          "content": "<p>They provide the weights as ckpts. I have not implemented this part yet (as I’m focused on training from scratch… the notebook and post will be shared within 2 hours). </p>\n<p>I assume you could build the model and load the weights by name maybe? I’m not sure though…</p>\n<p>That being said I will make it a priority to demonstrate that functionality… probably tomorrow.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1310670,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2021-05-16T20:15:14.137000",
          "content": "<p>I did get a small step further, the <code>EffNetV2Model</code> extends <code>tf.keras.Model</code> and does therefore support the <code>load_weights</code> functions. However, an error is thrown when the checkpoint is loaded this way.</p>\n<pre><code>m.load_weights('path_to_weights/efficientnetv2-s-21k-ft1k/model')\n</code></pre>\n<p>Output:</p>\n<pre><code>AssertionError: Some objects had attributes which were not restored:\n\n[A huge amount of layers which can't be loaded...]\n</code></pre>\n<p>The tutorial I followed for loading the checkpoint can be found <a href=\"https://www.tensorflow.org/tutorials/keras/save_and_load#checkpoint_callback_usage\" target=\"_blank\">here</a></p>\n<p>Happy to hear when you get any progress.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1310696,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-16T21:11:07.803000",
          "content": "<p>The weight in 4 files is the default TF format (not keras). Thus you need to use the default TF checkpoint loader from directory: <code>tf.train.Checkpoint</code>  </p>\n<pre><code>model = effnetv2_model.EffNetV2Model('efficientnetv2-s')\nCKPT_dir ='directory_of_dowloaded_weight'\ncheckpoint = tf.train.Checkpoint(model)\ncheckpoint.restore(tf.train.latest_checkpoint(CKPT_dir ))\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1310729,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-16T22:18:27.790000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1310744,
          "author_name": "Andrew Shao",
          "author_url": "",
          "post_date": "2021-05-16T22:53:01.850000",
          "content": "<p>Did anyone get this to work. I'm still just getting: AssertionError: Some objects had attributes which were not restored, when checking the loading status, and every weight is listed there(I.e No weight is loaded, but they all exist).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1310790,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-17T00:57:42.650000",
          "content": "<p>I'm making a notebook now… but this is from the infer.py file within the repo… seems straightforward…</p>\n<pre><code>def build_tf2_model():\n  \"\"\"Build the tf2 model.\"\"\"\n  # Use 'mixed_float16' if running on GPUs.\n  policy = tf.keras.mixed_precision.Policy('mixed_float16')\n  tf.keras.mixed_precision.set_global_policy(policy)\n  tf.config.run_functions_eagerly(FLAGS.debug)\n  # Create and run the model.\n  model = create_model(FLAGS.model_name, FLAGS.dataset_cfg, FLAGS.hparam_str)\n  # Use call (not build) to match the namescope: tensorflow issues/29576\n  model(tf.ones([1, 224, 224, 3]), False)\n  if FLAGS.model_dir:\n    ckpt = FLAGS.model_dir\n    if tf.io.gfile.isdir(ckpt):\n      ckpt = tf.train.latest_checkpoint(FLAGS.model_dir)\n    model.load_weights(ckpt)\n  model.summary()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1310796,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-17T01:08:04.013000",
          "content": "<p>Here… I made it into a notebook. I load the weights and infer on a random image of a cat.</p>\n<p><strong><a href=\"https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights\" target=\"_blank\">https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights</a></strong></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1310799,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-17T01:14:24.740000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1310968,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-17T04:55:34.250000",
          "content": "<p><a href=\"https://www.kaggle.com/andrewshao05\" target=\"_blank\">@andrewshao05</a> The snippet I give, works for me.  I am training right now with the pretrained 21k + finetune 1k weight. </p>\n<p>It may be need some cautious with your optimizer. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1310983,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-17T05:02:51.650000",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>   Does this approach work for finetuning (not inference) ? </p>\n<p>In general, I think TF checkpoint management is unnecessarily complicated and misleading with too many formats,  while in Pytorch everything is simple given you can save and load straightforwardly all states dict (last epoch, lr_schedule, model and optimizer weights) with the same single file. </p>\n<p>That's why I prefer to use plain numpy arrays and manage my own checkpoint during training and use neither TF checkpoint nor keras default format. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329232,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2021-05-31T01:40:35.640000",
          "content": "<p>InvalidArgumentError: Unsuccessful TensorSliceReader constructor: Failed to get matching files on /content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model: Unimplemented: File system scheme '[local]' not implemented (file: '/content/kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-b0/model')</p>\n<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> after put your codes why I get error like this? Can you please suggests me what was the problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329919,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2021-05-31T13:21:30.873000",
          "content": "<p>If you are using TPU's this error can occur when using dataset paths instead of Google Cloud paths. Make your dataset public and use the Google Cloud path to the weights.</p>\n<pre><code>GOOGLE_CLOUD_PATH= KaggleDatasets().get_gcs_path('your-dataset')\n\n# load model\nmodel= effnetv2_model.EffNetV2Model(model_name='efficientnetv2-b3')\n# dummy call to initialize model\nmodel(tf.ones((1,224,224,3)), training=False)\n\n# weight path on Google Cloud\nWEIGHT_PATH = f'{GOOGLE_CLOUD_PATH}/efficientnetv2-b3'\n# Get the latest checkpoint from path\nckpt = tf.train.latest_checkpoint(WEIGHT_PATH)\n\n# Load the weights\nmodel.load_weights(ckpt)\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1330778,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2021-06-01T05:07:32.297000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>.  Any suggestion for adding \"efficientnet-v2-s\" from \"efficientnetv2-s-21k\" I actually wrote \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s-21k\")\" but got error also I tried \"effnetv2_model.EffNetV2Model(name=\"efficientnetv2-s\")\" but still got error.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331351,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2021-06-01T12:16:01.303000",
          "content": "<p>For me using the following method works:</p>\n<p><code>model = effnetv2_model.EffNetV2Model(model_name='efficientnetv2-s')</code></p>\n<p>What kind of error do you get when loading the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331415,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2021-06-01T13:06:10.117000",
          "content": "<p>I found <strong>AssertionError</strong>:<br>\n&lt;AutoCastVariable 'efficientnetv2-s/stem/conv2d/kernel:0' shape=(3, 3, 3, 24) dtype=float32 dtype_to_cast_to=float32, numpy=<br>\narray([[[[ 1.31866440e-01,  2.92598866e-02,  5.63979149e-02,<br>\n           6.79737628e-02, -1.03196532e-01, -1.61898583e-02,<br>\n          -9.67781842e-02,  9.58853662e-02, -1.21669516e-01,<br>\n          -3.78544554e-02, -4.95341495e-02, -1.13189340e-01,<br>\n          -3.57197300e-02,  1.19056180e-01, -3.08847930e-02,<br>\n          -5.60378395e-02,  2.59503573e-02, -4.53319354e-03,<br>\n           3.76539188e-03, -2.62582731e-02, -4.90069948e-03,<br>\n          -2.31419038e-02,  1.32807806e-01,  1.46531919e-02],<br>\n         [ 2.10016251e-01, -5.64353578e-02,  2.44152751e-02,<br>\n           7.11440295e-02, -7.25007206e-02, -6.13373555e-02,<br>\n          -1.60715938e-01,  9.48695615e-02, -1.58229873e-01,<br>\n          -1.13155447e-01, -5.96799143e-02,  9.18575004e-02,<br>\n           2.68904697e-02, -2.05748565e-02, -1.37787834e-01,<br>\n          -4.67350520e-02, -7.78039470e-02,  7.96402842e-02,<br>\n           5.23473695e-02,  1.15151033e-01,  5.91931865e-03,<br>\n          -6.38899282e-02, -2.02910174e-02,  4.16407287e-02],<br>\n     ………………………………………………………………………………………….<br>\n     ………………………………………………………………………………………….</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331419,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2021-06-01T13:08:21.007000",
          "content": "<p>I used for dataset directory from kaggle : </p>\n<blockquote>\n  <p>gs://kds-5f97af08450b1cbaa872671ae6c64ff393755c658eeba01d3fb32b22/efficientnetv2-s-21k</p>\n</blockquote>\n<p>And when I try to load weights I found this error</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331439,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-06-01T13:19:57.007000",
          "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> - Looks like you are using the <strong><code>name</code></strong> argument instead of the <strong><code>model_name</code></strong>  argument. One is the string name associated with the keras model name (<strong><code>name</code></strong>) and the other informs on which backbone to load (<strong><code>model_name</code></strong>).</p>\n<p>Try changing that argument and seeing what happens?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331450,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2021-06-01T13:28:11.927000",
          "content": "<blockquote>\n  <p>ev2_s = effnetv2_model.EffNetV2Model(model_name=\"efficientnetv2-s\", name=\"efficientnetv2-s\") </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> I actually used your code and change Ii with \"efficientnetv2-s\" from \"efficientnetv2-b0\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332304,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-06-02T03:31:51.547000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> - this method works for fine tuning. I’ll be open sourcing a notebook showing the same in the COVID competition shortly and will tag you when it’s complete.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1308697": "Full code and pretrained weights for efficientnet v2 can be found here - https://github.com/google/automl/tree/master/efficientnetv2\n\n\n\n![Efficientnet Weights](https://github.com/google/automl/raw/master/efficientnetv2/g3doc/train_params.png)",
    "1310691": "End-to-end pipeline including training, inference on the test set, and submission using an EfficientNetV2 is now complete!! The whole thing takes less than 3 hours.\n* **Resulting score is 4.97 Levenshtein Distance on the public LB.**\n\n---\n\nHere is the [notebook](https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs)\n\nHere is my [discussion post](https://www.kaggle.com/c/bms-molecular-translation/discussion/239595) which links to all the relevant stuff.\n\n---\n\nSide Note: Here is a [notebook](https://www.kaggle.com/dschettler8845/load-efficientnetv2-pretrained-weights) I created showing how to load the pre-trained weights. I do inference on a random photo to show that it works.\n\n---\n\nHope this helps!",
    "1309407": "Will implement and share a notebook leveraging keras and tpu this weekend. **[UPDATE: DONE]**",
    "1309746": "It seems the default weight is the pretrained Imagenet1k \n\nHowever the bigger Imagenet (21k + finetune on 1k) give better results on the paper.   I'm experimenting with this last one to see (a bit trickier to load the weight in the context of distributed training)",
    "1312980": "Based on my first expriments,  EffnetV1  with noisy student weight give better results in this dataset than EFFnetV2.\n\nI think EFFnetV2 would benefit from NS training given how it improved significantly EffnetV1. \n\nMay be the comparaison in the paper is with EffnetV1 without NS training. ",
    "1310637": "Happy to hear the models and weights are finally made available.\nI did take a look at it, however there does not seem to be an option to load the pretrained weights when building the model.\n\n```\n# build keras model\nmodel = effnetv2_model.EffNetV2Model('efficientnetv2-s')\n# how to load the pretrained weights?\n...\n```\n\nWhen unzipping the checkpoints there are 4 files:\n\n```\ncheckpoint\nmodel.data-00000-of-00001\nmodel.index\nmodel.meta\n```\n\nConventionally the weights are given as a single h5 file, how to load the given weights?"
  }
}