{
  "id": 130378,
  "title": "Can't load weights for NASNetLarge model",
  "url": "/competitions/flower-classification-with-tpus/discussion/130378",
  "author_name": "",
  "post_date": "2020-02-13T19:28:15.446556500Z",
  "votes": 3,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I'd like to try NASNetLarge model. Unfortunately, it's <code>include_top</code> option is bugged. It can be fixed with manual loading of NASNet-large-no-top.h5 file with <code>load_weights</code> function. My problem is that I can't find a way to do it. I tried to load h5 file to GCS with <code>get_gcs_path</code>, but I'm getting \"Unable to load file\" error all the time when I apply <code>load_weights</code> function. Is it possible at all? I use NASNet-large-no-top public dataset added to my notebook.</p>",
  "messages": [
    {
      "id": "745398",
      "postDate": "02/13/2020 19:28:15",
      "content": "<p>I'd like to try NASNetLarge model. Unfortunately, it's <code>include_top</code> option is bugged. It can be fixed with manual loading of NASNet-large-no-top.h5 file with <code>load_weights</code> function. My problem is that I can't find a way to do it. I tried to load h5 file to GCS with <code>get_gcs_path</code>, but I'm getting \"Unable to load file\" error all the time when I apply <code>load_weights</code> function. Is it possible at all? I use NASNet-large-no-top public dataset added to my notebook.</p>",
      "rawMarkdown": "I'd like to try NASNetLarge model. Unfortunately, it's `include_top` option is bugged. It can be fixed with manual loading of NASNet-large-no-top.h5 file with `load_weights` function. My problem is that I can't find a way to do it. I tried to load h5 file to GCS with `get_gcs_path`, but I'm getting \"Unable to load file\" error all the time when I apply `load_weights` function. Is it possible at all? I use NASNet-large-no-top public dataset added to my notebook.",
      "votes": null
    },
    {
      "id": "745403",
      "postDate": "02/13/2020 19:33:42",
      "content": "<p>The proper way of adding external data to your TPU model is to upload the data as a public Kaggle Dataset, attach the dataset to your notebook, then use\n<code>\ngcs_path = KaggleDatasets().get_gcs_path('dataset_directory_name')\n</code></p>",
      "rawMarkdown": "The proper way of adding external data to your TPU model is to upload the data as a public Kaggle Dataset, attach the dataset to your notebook, then use\n```\ngcs_path = KaggleDatasets().get_gcs_path('dataset_directory_name')\n```",
      "votes": null
    },
    {
      "id": "745405",
      "postDate": "02/13/2020 19:36:44",
      "content": "<p>I've done the following (I have public nasnetlargenotop dataset added to my notebook):\n<code>GCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('nasnetlargenotop')</code>\n...\n<code>pretrained_model.load_weights(GCS_DS_PATH_2 +'\\NASNet-large-no-top.h5')</code></p>\n\n<p>But I get the following error:\nOSError: Unable to open file (unable to open file: name = 'gs://kds-486053fd6a5e92cda576498de5377befe2d1c935fb4e8de1c252afb2/NASNet-large-no-top.h5', errno = 2, error message = 'No such file or directory', flags = 0, o_flags = 0)</p>\n\n<p>(I omitted code non-relevant to the problem. Of course I have the model loaded.)</p>",
      "rawMarkdown": "I've done the following (I have public nasnetlargenotop dataset added to my notebook):\n`GCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('nasnetlargenotop')`\n...\n`pretrained_model.load_weights(GCS_DS_PATH_2 +'\\NASNet-large-no-top.h5')`\n\nBut I get the following error:\nOSError: Unable to open file (unable to open file: name = 'gs://kds-486053fd6a5e92cda576498de5377befe2d1c935fb4e8de1c252afb2/NASNet-large-no-top.h5', errno = 2, error message = 'No such file or directory', flags = 0, o_flags = 0)\n\n(I omitted code non-relevant to the problem. Of course I have the model loaded.)",
      "votes": null
    },
    {
      "id": "745409",
      "postDate": "02/13/2020 19:40:12",
      "content": "<p>You can use gsutil on the command line to troubleshoot. Maybe the path is not right.\n<code>\n!gsutil ls gs://...\n</code></p>",
      "rawMarkdown": "You can use gsutil on the command line to troubleshoot. Maybe the path is not right.\n```\n!gsutil ls gs://...\n```",
      "votes": null
    },
    {
      "id": "745411",
      "postDate": "02/13/2020 19:43:53",
      "content": "<p>I tried gsutil, The path ckecks out ... I see you file there. 🤔</p>",
      "rawMarkdown": "I tried gsutil, The path ckecks out ... I see you file there. 🤔",
      "votes": null
    },
    {
      "id": "745413",
      "postDate": "02/13/2020 19:45:59",
      "content": "<p>Hmm, I seem to remember the .h5 loader in Keras does not like GCS... ouch.</p>",
      "rawMarkdown": "Hmm, I seem to remember the .h5 loader in Keras does not like GCS... ouch.",
      "votes": null
    },
    {
      "id": "745415",
      "postDate": "02/13/2020 19:49:53",
      "content": "<p>You'll have to do some juggling here. First, put your model creation in a create_model() function so that you can call it multiple times. Then:</p>\n\n<p>```\n1) with strategy.scope():\n              tpu_model = create_model()\n2) cpu_model =  create_model()\n3) cpu_model.load_weights('local path, NOT GCS')\n4) tpu_model.set_weights(cpu_model.get_weights())</p>\n\n<p>```\nI have not tried myself yet. Bugs possible.</p>",
      "rawMarkdown": "You'll have to do some juggling here. First, put your model creation in a create_model() function so that you can call it multiple times. Then:\n\n```\n1) with strategy.scope():\n              tpu_model = create_model()\n2) cpu_model =  create_model()\n3) cpu_model.load_weights('local path, NOT GCS')\n4) tpu_model.set_weights(cpu_model.get_weights())\n\n```\nI have not tried myself yet. Bugs possible.",
      "votes": null
    },
    {
      "id": "745417",
      "postDate": "02/13/2020 19:50:49",
      "content": "<p>OK, I'll try.</p>",
      "rawMarkdown": "OK, I'll try.",
      "votes": null
    },
    {
      "id": "745443",
      "postDate": "02/13/2020 20:19:08",
      "content": "<p>Unfortunately, this doesn't work. It gives very long error log (<a href=\"https://pastebin.com/WKZA8nZ3\">full log</a>) with the following message in the end: \"Make sure the slot variables are created under the same strategy scope. This may happen if you're restoring from a checkpoint outside the scope\".</p>",
      "rawMarkdown": "Unfortunately, this doesn't work. It gives very long error log ([full log](https://pastebin.com/WKZA8nZ3)) with the following message in the end: \"Make sure the slot variables are created under the same strategy scope. This may happen if you're restoring from a checkpoint outside the scope\".",
      "votes": null
    },
    {
      "id": "745463",
      "postDate": "02/13/2020 20:52:05",
      "content": "<p>That problem does not sound too scary. You probably just need to remove the .compile code from create_model() and do that at the end, after the juggling.</p>",
      "rawMarkdown": "That problem does not sound too scary. You probably just need to remove the .compile code from create_model() and do that at the end, after the juggling.",
      "votes": null
    },
    {
      "id": "745464",
      "postDate": "02/13/2020 20:56:06",
      "content": "<p>In fact I called .compile after all juggling, so this is not the case.</p>",
      "rawMarkdown": "In fact I called .compile after all juggling, so this is not the case.",
      "votes": null
    },
    {
      "id": "745490",
      "postDate": "02/13/2020 21:31:09",
      "content": "<p>Was the .compile in or out of the strategy.scope()? It should not matter, just asking.\nAlso, is the .compile on the correct model ?</p>\n\n<p>If there's nothing obvious like that then, hmm, I'll have to try.</p>",
      "rawMarkdown": "Was the .compile in or out of the strategy.scope()? It should not matter, just asking.\nAlso, is the .compile on the correct model ?\n\nIf there's nothing obvious like that then, hmm, I'll have to try.",
      "votes": null
    },
    {
      "id": "745503",
      "postDate": "02/13/2020 21:40:24",
      "content": "<p>.compile is out of strategy.scope()</p>\n\n<p>And this is a model I compile:\n```\nmodel = tf.keras.Sequential([\n    tpu_model,\n    tf.keras.layers.GlobalAveragePooling2D(),\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax')\n])</p>\n\n<p>```</p>",
      "rawMarkdown": ".compile is out of strategy.scope()\n\nAnd this is a model I compile:\n```\nmodel = tf.keras.Sequential([\n    tpu_model,\n    tf.keras.layers.GlobalAveragePooling2D(),\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax')\n])\n\n```",
      "votes": null
    },
    {
      "id": "745506",
      "postDate": "02/13/2020 21:46:46",
      "content": "<p>Try the .compile in the <code>strategy.scope()</code>. The model is supposed to remember where it was created but maybe there is a bug.</p>\n\n<p>The general rule is that all variable creations must be in the strategy.scope() if these variables are supposed to be on the TPU. So definitely all Keras layers myst be defined in scope. .compile defines variables too so theoretically it should be in the scope as well. We added some code in .compile() to set up the scope automatically if the model was created in the scope, for the sake of people who forget. Maybe it's not working as well as it should.</p>\n\n<p>Please also check that <code>tf.keras.layers.Dense</code> in your code snippet above is created in the <code>strategy.scope()</code>.</p>",
      "rawMarkdown": "Try the .compile in the `strategy.scope()`. The model is supposed to remember where it was created but maybe there is a bug.\n\nThe general rule is that all variable creations must be in the strategy.scope() if these variables are supposed to be on the TPU. So definitely all Keras layers myst be defined in scope. .compile defines variables too so theoretically it should be in the scope as well. We added some code in .compile() to set up the scope automatically if the model was created in the scope, for the sake of people who forget. Maybe it's not working as well as it should.\n\nPlease also check that `tf.keras.layers.Dense` in your code snippet above is created in the `strategy.scope()`.",
      "votes": null
    },
    {
      "id": "745520",
      "postDate": "02/13/2020 22:03:03",
      "content": "<p>It seems to be working, because I got now ResourceExhaustedError: {{function_node __inference_distributed_function_608788}} Compilation failure: Ran out of memory in memory space hbm. Used 16.13G of 16.00G hbm. Exceeded hbm capacity by 133.58M. </p>\n\n<p>Seems that NASNetLarge is too large for Kaggle environment.😃 </p>",
      "rawMarkdown": "It seems to be working, because I got now ResourceExhaustedError: {{function_node __inference_distributed_function_608788}} Compilation failure: Ran out of memory in memory space hbm. Used 16.13G of 16.00G hbm. Exceeded hbm capacity by 133.58M. \n\nSeems that NASNetLarge is too large for Kaggle environment.😃",
      "votes": null
    },
    {
      "id": "745524",
      "postDate": "02/13/2020 22:10:05",
      "content": "<p>\\o/ yay, good news, can you share the exact juggling code for others ?</p>\n\n<p>For you memory problem: reduce batch size, reduce image size. Probably not by much. It looks like you're just a hair above.</p>",
      "rawMarkdown": "\\o/ yay, good news, can you share the exact juggling code for others ?\n\nFor you memory problem: reduce batch size, reduce image size. Probably not by much. It looks like you're just a hair above.",
      "votes": null
    },
    {
      "id": "745881",
      "postDate": "02/14/2020 10:10:21",
      "content": "<p>I decreased image size to 331, and it works now! Thank you for your help!👍 </p>",
      "rawMarkdown": "I decreased image size to 331, and it works now! Thank you for your help!👍",
      "votes": null
    },
    {
      "id": "745926",
      "postDate": "02/14/2020 11:45:22",
      "content": "<p>Also this problem can be solved by decreasing BATCH_SIZE to 8 (while keeping image size equal to 512).</p>",
      "rawMarkdown": "Also this problem can be solved by decreasing BATCH_SIZE to 8 (while keeping image size equal to 512).",
      "votes": null
    },
    {
      "id": "746156",
      "postDate": "02/14/2020 17:23:19",
      "content": "<p>Batch size 8 or 8*8=64 ?</p>",
      "rawMarkdown": "Batch size 8 or 8*8=64 ?",
      "votes": null
    },
    {
      "id": "746157",
      "postDate": "02/14/2020 17:26:03",
      "content": "<p>Also, would you mind publishing the code showing exactly what you did. It's unfortunate a workaround is needed here but I'm sure other people would be interested. You don't have to publish your leaderboard submission. The <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">getting started notebook</a> with the NASNet modification will do.</p>",
      "rawMarkdown": "Also, would you mind publishing the code showing exactly what you did. It's unfortunate a workaround is needed here but I'm sure other people would be interested. You don't have to publish your leaderboard submission. The [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/) with the NASNet modification will do.",
      "votes": null
    },
    {
      "id": "746475",
      "postDate": "02/15/2020 04:14:24",
      "content": "<p>You can use NASNet -\nfrom tensorflow.keras.applications import nasnet\nIMAGE_SIZE =[331, 331] # NASNetLarge</p>\n\n<p>with strategy.scope():\n    nnet = nasnet.NASNetLarge(\n        input_shape=(IMAGE_SIZE2[0], IMAGE_SIZE2[1], 3),\n        weights='imagenet',\n        include_top=False\n    )</p>",
      "rawMarkdown": "You can use NASNet -\nfrom tensorflow.keras.applications import nasnet\nIMAGE_SIZE =[331, 331] # NASNetLarge\n\nwith strategy.scope():\n    nnet = nasnet.NASNetLarge(\n        input_shape=(IMAGE_SIZE2[0], IMAGE_SIZE2[1], 3),\n        weights='imagenet',\n        include_top=False\n    )",
      "votes": null
    },
    {
      "id": "747017",
      "postDate": "02/15/2020 21:26:39",
      "content": "<p>Seems like a good idea. I'll do it.</p>",
      "rawMarkdown": "Seems like a good idea. I'll do it.",
      "votes": null
    },
    {
      "id": "747027",
      "postDate": "02/15/2020 21:37:37",
      "content": "<p>And about batch size - I meant 8*8.</p>",
      "rawMarkdown": "And about batch size - I meant 8*8.",
      "votes": null
    },
    {
      "id": "747370",
      "postDate": "02/16/2020 10:50:47",
      "content": "<p><a href=\"https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge\">Kernel with NASNetLarge</a></p>",
      "rawMarkdown": "[Kernel with NASNetLarge](https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge)",
      "votes": null
    },
    {
      "id": "749603",
      "postDate": "02/18/2020 19:53:49",
      "content": "<p>Araik was saying there was some bug with NASNetLarge and includetop=False ? Did you encounter it ?</p>",
      "rawMarkdown": "Araik was saying there was some bug with NASNetLarge and includetop=False ? Did you encounter it ?",
      "votes": null
    },
    {
      "id": "750035",
      "postDate": "02/19/2020 04:44:23",
      "content": "<p>No problem with using the code snippet I posted above.  It was keeping to the NASNet size [331,331] whereas Araik is using [512,512] per the notebook. Thought perhaps wanted to use own weights?  I used imagenet.  I also have latest available Docker selected maybe that makes a difference - github had EffficientNet added in Dec 2019 and was trying to import from tensorflow.keras.applications but seems not yet in docker versions so still need to do pip install for that. </p>\n\n<p>Still trialling things with NASNet, it looked promising - accuracy 1.000 but not so great yet on LB.  If any breakthroughs will publish a notebook. (Needs more sweet peas and red ginger!)</p>",
      "rawMarkdown": "No problem with using the code snippet I posted above.  It was keeping to the NASNet size [331,331] whereas Araik is using [512,512] per the notebook. Thought perhaps wanted to use own weights?  I used imagenet.  I also have latest available Docker selected maybe that makes a difference - github had EffficientNet added in Dec 2019 and was trying to import from tensorflow.keras.applications but seems not yet in docker versions so still need to do pip install for that. \n\nStill trialling things with NASNet, it looked promising - accuracy 1.000 but not so great yet on LB.  If any breakthroughs will publish a notebook. (Needs more sweet peas and red ginger!)",
      "votes": null
    },
    {
      "id": "768385",
      "postDate": "03/10/2020 18:19:25",
      "content": "<p>Thank you for making it public.</p>",
      "rawMarkdown": "Thank you for making it public.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 745403,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/13/2020 19:33:42",
      "content": "<p>The proper way of adding external data to your TPU model is to upload the data as a public Kaggle Dataset, attach the dataset to your notebook, then use\n<code>\ngcs_path = KaggleDatasets().get_gcs_path('dataset_directory_name')\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 745405,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 19:36:44",
          "content": "<p>I've done the following (I have public nasnetlargenotop dataset added to my notebook):\n<code>GCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('nasnetlargenotop')</code>\n...\n<code>pretrained_model.load_weights(GCS_DS_PATH_2 +'\\NASNet-large-no-top.h5')</code></p>\n\n<p>But I get the following error:\nOSError: Unable to open file (unable to open file: name = 'gs://kds-486053fd6a5e92cda576498de5377befe2d1c935fb4e8de1c252afb2/NASNet-large-no-top.h5', errno = 2, error message = 'No such file or directory', flags = 0, o_flags = 0)</p>\n\n<p>(I omitted code non-relevant to the problem. Of course I have the model loaded.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745409,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 19:40:12",
          "content": "<p>You can use gsutil on the command line to troubleshoot. Maybe the path is not right.\n<code>\n!gsutil ls gs://...\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745411,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 19:43:53",
          "content": "<p>I tried gsutil, The path ckecks out ... I see you file there. 🤔</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745413,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 19:45:59",
          "content": "<p>Hmm, I seem to remember the .h5 loader in Keras does not like GCS... ouch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745415,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 19:49:53",
          "content": "<p>You'll have to do some juggling here. First, put your model creation in a create_model() function so that you can call it multiple times. Then:</p>\n\n<p>```\n1) with strategy.scope():\n              tpu_model = create_model()\n2) cpu_model =  create_model()\n3) cpu_model.load_weights('local path, NOT GCS')\n4) tpu_model.set_weights(cpu_model.get_weights())</p>\n\n<p>```\nI have not tried myself yet. Bugs possible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745417,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 19:50:49",
          "content": "<p>OK, I'll try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745443,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 20:19:08",
          "content": "<p>Unfortunately, this doesn't work. It gives very long error log (<a href=\"https://pastebin.com/WKZA8nZ3\">full log</a>) with the following message in the end: \"Make sure the slot variables are created under the same strategy scope. This may happen if you're restoring from a checkpoint outside the scope\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745463,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 20:52:05",
          "content": "<p>That problem does not sound too scary. You probably just need to remove the .compile code from create_model() and do that at the end, after the juggling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745464,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 20:56:06",
          "content": "<p>In fact I called .compile after all juggling, so this is not the case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745490,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 21:31:09",
          "content": "<p>Was the .compile in or out of the strategy.scope()? It should not matter, just asking.\nAlso, is the .compile on the correct model ?</p>\n\n<p>If there's nothing obvious like that then, hmm, I'll have to try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745503,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 21:40:24",
          "content": "<p>.compile is out of strategy.scope()</p>\n\n<p>And this is a model I compile:\n```\nmodel = tf.keras.Sequential([\n    tpu_model,\n    tf.keras.layers.GlobalAveragePooling2D(),\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax')\n])</p>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745506,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 21:46:46",
          "content": "<p>Try the .compile in the <code>strategy.scope()</code>. The model is supposed to remember where it was created but maybe there is a bug.</p>\n\n<p>The general rule is that all variable creations must be in the strategy.scope() if these variables are supposed to be on the TPU. So definitely all Keras layers myst be defined in scope. .compile defines variables too so theoretically it should be in the scope as well. We added some code in .compile() to set up the scope automatically if the model was created in the scope, for the sake of people who forget. Maybe it's not working as well as it should.</p>\n\n<p>Please also check that <code>tf.keras.layers.Dense</code> in your code snippet above is created in the <code>strategy.scope()</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745520,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/13/2020 22:03:03",
          "content": "<p>It seems to be working, because I got now ResourceExhaustedError: {{function_node __inference_distributed_function_608788}} Compilation failure: Ran out of memory in memory space hbm. Used 16.13G of 16.00G hbm. Exceeded hbm capacity by 133.58M. </p>\n\n<p>Seems that NASNetLarge is too large for Kaggle environment.😃 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745524,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 22:10:05",
          "content": "<p>\\o/ yay, good news, can you share the exact juggling code for others ?</p>\n\n<p>For you memory problem: reduce batch size, reduce image size. Probably not by much. It looks like you're just a hair above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745881,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/14/2020 10:10:21",
          "content": "<p>I decreased image size to 331, and it works now! Thank you for your help!👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745926,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/14/2020 11:45:22",
          "content": "<p>Also this problem can be solved by decreasing BATCH_SIZE to 8 (while keeping image size equal to 512).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746156,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:23:19",
          "content": "<p>Batch size 8 or 8*8=64 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746157,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:26:03",
          "content": "<p>Also, would you mind publishing the code showing exactly what you did. It's unfortunate a workaround is needed here but I'm sure other people would be interested. You don't have to publish your leaderboard submission. The <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">getting started notebook</a> with the NASNet modification will do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747017,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/15/2020 21:26:39",
          "content": "<p>Seems like a good idea. I'll do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747027,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/15/2020 21:37:37",
          "content": "<p>And about batch size - I meant 8*8.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747370,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/16/2020 10:50:47",
          "content": "<p><a href=\"https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge\">Kernel with NASNetLarge</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768385,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/10/2020 18:19:25",
          "content": "<p>Thank you for making it public.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746475,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "02/15/2020 04:14:24",
      "content": "<p>You can use NASNet -\nfrom tensorflow.keras.applications import nasnet\nIMAGE_SIZE =[331, 331] # NASNetLarge</p>\n\n<p>with strategy.scope():\n    nnet = nasnet.NASNetLarge(\n        input_shape=(IMAGE_SIZE2[0], IMAGE_SIZE2[1], 3),\n        weights='imagenet',\n        include_top=False\n    )</p>",
      "votes": null,
      "replies": [
        {
          "id": 749603,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 19:53:49",
          "content": "<p>Araik was saying there was some bug with NASNetLarge and includetop=False ? Did you encounter it ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750035,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "02/19/2020 04:44:23",
          "content": "<p>No problem with using the code snippet I posted above.  It was keeping to the NASNet size [331,331] whereas Araik is using [512,512] per the notebook. Thought perhaps wanted to use own weights?  I used imagenet.  I also have latest available Docker selected maybe that makes a difference - github had EffficientNet added in Dec 2019 and was trying to import from tensorflow.keras.applications but seems not yet in docker versions so still need to do pip install for that. </p>\n\n<p>Still trialling things with NASNet, it looked promising - accuracy 1.000 but not so great yet on LB.  If any breakthroughs will publish a notebook. (Needs more sweet peas and red ginger!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "745398": "I'd like to try NASNetLarge model. Unfortunately, it's `include_top` option is bugged. It can be fixed with manual loading of NASNet-large-no-top.h5 file with `load_weights` function. My problem is that I can't find a way to do it. I tried to load h5 file to GCS with `get_gcs_path`, but I'm getting \"Unable to load file\" error all the time when I apply `load_weights` function. Is it possible at all? I use NASNet-large-no-top public dataset added to my notebook.",
    "745403": "The proper way of adding external data to your TPU model is to upload the data as a public Kaggle Dataset, attach the dataset to your notebook, then use\n```\ngcs_path = KaggleDatasets().get_gcs_path('dataset_directory_name')\n```",
    "745405": "I've done the following (I have public nasnetlargenotop dataset added to my notebook):\n`GCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('nasnetlargenotop')`\n...\n`pretrained_model.load_weights(GCS_DS_PATH_2 +'\\NASNet-large-no-top.h5')`\n\nBut I get the following error:\nOSError: Unable to open file (unable to open file: name = 'gs://kds-486053fd6a5e92cda576498de5377befe2d1c935fb4e8de1c252afb2/NASNet-large-no-top.h5', errno = 2, error message = 'No such file or directory', flags = 0, o_flags = 0)\n\n(I omitted code non-relevant to the problem. Of course I have the model loaded.)",
    "745409": "You can use gsutil on the command line to troubleshoot. Maybe the path is not right.\n```\n!gsutil ls gs://...\n```",
    "745411": "I tried gsutil, The path ckecks out ... I see you file there. 🤔",
    "745413": "Hmm, I seem to remember the .h5 loader in Keras does not like GCS... ouch.",
    "745415": "You'll have to do some juggling here. First, put your model creation in a create_model() function so that you can call it multiple times. Then:\n\n```\n1) with strategy.scope():\n              tpu_model = create_model()\n2) cpu_model =  create_model()\n3) cpu_model.load_weights('local path, NOT GCS')\n4) tpu_model.set_weights(cpu_model.get_weights())\n\n```\nI have not tried myself yet. Bugs possible.",
    "745417": "OK, I'll try.",
    "745443": "Unfortunately, this doesn't work. It gives very long error log ([full log](https://pastebin.com/WKZA8nZ3)) with the following message in the end: \"Make sure the slot variables are created under the same strategy scope. This may happen if you're restoring from a checkpoint outside the scope\".",
    "745463": "That problem does not sound too scary. You probably just need to remove the .compile code from create_model() and do that at the end, after the juggling.",
    "745464": "In fact I called .compile after all juggling, so this is not the case.",
    "745490": "Was the .compile in or out of the strategy.scope()? It should not matter, just asking.\nAlso, is the .compile on the correct model ?\n\nIf there's nothing obvious like that then, hmm, I'll have to try.",
    "745503": ".compile is out of strategy.scope()\n\nAnd this is a model I compile:\n```\nmodel = tf.keras.Sequential([\n    tpu_model,\n    tf.keras.layers.GlobalAveragePooling2D(),\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax')\n])\n\n```",
    "745506": "Try the .compile in the `strategy.scope()`. The model is supposed to remember where it was created but maybe there is a bug.\n\nThe general rule is that all variable creations must be in the strategy.scope() if these variables are supposed to be on the TPU. So definitely all Keras layers myst be defined in scope. .compile defines variables too so theoretically it should be in the scope as well. We added some code in .compile() to set up the scope automatically if the model was created in the scope, for the sake of people who forget. Maybe it's not working as well as it should.\n\nPlease also check that `tf.keras.layers.Dense` in your code snippet above is created in the `strategy.scope()`.",
    "745520": "It seems to be working, because I got now ResourceExhaustedError: {{function_node __inference_distributed_function_608788}} Compilation failure: Ran out of memory in memory space hbm. Used 16.13G of 16.00G hbm. Exceeded hbm capacity by 133.58M. \n\nSeems that NASNetLarge is too large for Kaggle environment.😃",
    "745524": "\\o/ yay, good news, can you share the exact juggling code for others ?\n\nFor you memory problem: reduce batch size, reduce image size. Probably not by much. It looks like you're just a hair above.",
    "745881": "I decreased image size to 331, and it works now! Thank you for your help!👍",
    "745926": "Also this problem can be solved by decreasing BATCH_SIZE to 8 (while keeping image size equal to 512).",
    "746156": "Batch size 8 or 8*8=64 ?",
    "746157": "Also, would you mind publishing the code showing exactly what you did. It's unfortunate a workaround is needed here but I'm sure other people would be interested. You don't have to publish your leaderboard submission. The [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/) with the NASNet modification will do.",
    "746475": "You can use NASNet -\nfrom tensorflow.keras.applications import nasnet\nIMAGE_SIZE =[331, 331] # NASNetLarge\n\nwith strategy.scope():\n    nnet = nasnet.NASNetLarge(\n        input_shape=(IMAGE_SIZE2[0], IMAGE_SIZE2[1], 3),\n        weights='imagenet',\n        include_top=False\n    )",
    "747017": "Seems like a good idea. I'll do it.",
    "747027": "And about batch size - I meant 8*8.",
    "747370": "[Kernel with NASNetLarge](https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge)",
    "749603": "Araik was saying there was some bug with NASNetLarge and includetop=False ? Did you encounter it ?",
    "750035": "No problem with using the code snippet I posted above.  It was keeping to the NASNet size [331,331] whereas Araik is using [512,512] per the notebook. Thought perhaps wanted to use own weights?  I used imagenet.  I also have latest available Docker selected maybe that makes a difference - github had EffficientNet added in Dec 2019 and was trying to import from tensorflow.keras.applications but seems not yet in docker versions so still need to do pip install for that. \n\nStill trialling things with NASNet, it looked promising - accuracy 1.000 but not so great yet on LB.  If any breakthroughs will publish a notebook. (Needs more sweet peas and red ginger!)",
    "768385": "Thank you for making it public."
  },
  "source": "meta"
}