{
  "id": 399359,
  "title": "TPU VM 3-8 suddenly not working?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/399359",
  "author_name": "",
  "post_date": "2023-04-03T18:25:43.818115400Z",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Not sure if anybody else is experiencing issues with the TPU VM accelerators? My training notebook that used to work on Saturday now seems to be not using the TPU…..the speed achieved looks more like a very descent desktop ;-)</p>\n<p>I get the errors below. I tried with the default TF 2.11 installed and also with the install of TF 2.9.1…both give the same errors.</p>\n<p>With only 20 hours of TPU time quota there is not a lot of room for troubleshooting.</p>\n<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Mentioning you because I'am not sure if Kaggle knows this error?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2Fee602c4fd6b90bd35362b80aedb5ae7d%2Ftpu.png?generation=1680546335585738&amp;alt=media\" alt=\"\"></p>\n<p>Update: The same error happens on the Example notebook that is provided on the Product Feedback page:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F2e92b7ba14ca0dea57eefbe9833b9ec1%2Ftpu2.png?generation=1680546901060758&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2207991",
      "postDate": "04/03/2023 18:25:43",
      "content": "<p>Not sure if anybody else is experiencing issues with the TPU VM accelerators? My training notebook that used to work on Saturday now seems to be not using the TPU…..the speed achieved looks more like a very descent desktop ;-)</p>\n<p>I get the errors below. I tried with the default TF 2.11 installed and also with the install of TF 2.9.1…both give the same errors.</p>\n<p>With only 20 hours of TPU time quota there is not a lot of room for troubleshooting.</p>\n<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Mentioning you because I'am not sure if Kaggle knows this error?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2Fee602c4fd6b90bd35362b80aedb5ae7d%2Ftpu.png?generation=1680546335585738&amp;alt=media\" alt=\"\"></p>\n<p>Update: The same error happens on the Example notebook that is provided on the Product Feedback page:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F2e92b7ba14ca0dea57eefbe9833b9ec1%2Ftpu2.png?generation=1680546901060758&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Not sure if anybody else is experiencing issues with the TPU VM accelerators? My training notebook that used to work on Saturday now seems to be not using the TPU.....the speed achieved looks more like a very descent desktop ;-)\n\nI get the errors below. I tried with the default TF 2.11 installed and also with the install of TF 2.9.1...both give the same errors.\n\nWith only 20 hours of TPU time quota there is not a lot of room for troubleshooting.\n\n@herbison Mentioning you because I'am not sure if Kaggle knows this error?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2Fee602c4fd6b90bd35362b80aedb5ae7d%2Ftpu.png?generation=1680546335585738&alt=media)\n\nUpdate: The same error happens on the Example notebook that is provided on the Product Feedback page:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F2e92b7ba14ca0dea57eefbe9833b9ec1%2Ftpu2.png?generation=1680546901060758&alt=media)",
      "votes": null
    },
    {
      "id": "2208039",
      "postDate": "04/03/2023 18:51:45",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> The \"latest\" tf (v130 March 29 2023) was rolled back because of bug with juptyer-lsp locking up the entire notebook for TPU VMs. This was the first build of the TPU VM that came with tf for TPU VM pre-installed (vs. needing to be installed).</p>\n<p>I have a fixed version (v131 Apr 3 2023) which is unreleased but seems to fix the lsp issue and otherwise be identical to the v130 release. Hoping for that to be released early this week.</p>\n<p>In the meantime, you'll need to pip install the tf tpu vm package to get TPU VM speedup.</p>",
      "rawMarkdown": "rsmits The \"latest\" tf (v130 March 29 2023) was rolled back because of bug with juptyer-lsp locking up the entire notebook for TPU VMs. This was the first build of the TPU VM that came with tf for TPU VM pre-installed (vs. needing to be installed).\n\nI have a fixed version (v131 Apr 3 2023) which is unreleased but seems to fix the lsp issue and otherwise be identical to the v130 release. Hoping for that to be released early this week.\n\nIn the meantime, you'll need to pip install the tf tpu vm package to get TPU VM speedup.",
      "votes": null
    },
    {
      "id": "2208053",
      "postDate": "04/03/2023 19:01:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Thanks for the quick reply. Ok I will keep the above in mind.</p>\n<p>With the pip install for the tf tpu vm package you mean uncomment the line that installs the TF 2.9.1 version?</p>\n<p>Or keep the default TF 2.11 and install what specific package?</p>",
      "rawMarkdown": "Hi @herbison Thanks for the quick reply. Ok I will keep the above in mind.\n\nWith the pip install for the tf tpu vm package you mean uncomment the line that installs the TF 2.9.1 version?\n\nOr keep the default TF 2.11 and install what specific package?",
      "votes": null
    },
    {
      "id": "2208061",
      "postDate": "04/03/2023 19:06:26",
      "content": "<p>It would be uncomment to install tf 2.9.1, you can also install other tpu vm versions of tensorflow, but you'll need to update the libtpu version as well, example for tf 2.11:<br>\n<a href=\"https://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm\" target=\"_blank\">https://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm</a></p>\n<p>There's a table of versions:<br>\n<a href=\"https://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm\" target=\"_blank\">https://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm</a></p>",
      "rawMarkdown": "It would be uncomment to install tf 2.9.1, you can also install other tpu vm versions of tensorflow, but you'll need to update the libtpu version as well, example for tf 2.11:\nhttps://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm\n\nThere's a table of versions:\nhttps://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm",
      "votes": null
    },
    {
      "id": "2208090",
      "postDate": "04/03/2023 19:43:10",
      "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Ok tried both of the above…with TF 2.9.1 and the lines for installing TF 2.11. With some delays .. the notebook gets to the point where it would start training. Then show different kinds of 'model_pruner' errors and then seems to hang or reboot the kernel.</p>\n<p>I think I'll wait for the updated notebook version to be released as this is killing my TPU quota. Thanks for the support anyway.</p>",
      "rawMarkdown": "herbison Ok tried both of the above...with TF 2.9.1 and the lines for installing TF 2.11. With some delays .. the notebook gets to the point where it would start training. Then show different kinds of 'model_pruner' errors and then seems to hang or reboot the kernel.\n\nI think I'll wait for the updated notebook version to be released as this is killing my TPU quota. Thanks for the support anyway.",
      "votes": null
    },
    {
      "id": "2208092",
      "postDate": "04/03/2023 19:44:31",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Feel free to share the notebook with me too and I can evaluate it.</p>",
      "rawMarkdown": "rsmits Feel free to share the notebook with me too and I can evaluate it.",
      "votes": null
    },
    {
      "id": "2208116",
      "postDate": "04/03/2023 19:57:50",
      "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> ok send you a direct message with some questions regarding that.</p>",
      "rawMarkdown": "herbison ok send you a direct message with some questions regarding that.",
      "votes": null
    },
    {
      "id": "2210668",
      "postDate": "04/05/2023 14:46:28",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> The \"latest\" TPU VM image should have tf2.12 for TPU VM preinstalled correctly.</p>",
      "rawMarkdown": "rsmits The \"latest\" TPU VM image should have tf2.12 for TPU VM preinstalled correctly.",
      "votes": null
    },
    {
      "id": "2210733",
      "postDate": "04/05/2023 15:32:30",
      "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Ok thanks for giving a notice. I will try later today and let you know the results !</p>",
      "rawMarkdown": "herbison Ok thanks for giving a notice. I will try later today and let you know the results !",
      "votes": null
    },
    {
      "id": "2211100",
      "postDate": "04/05/2023 19:48:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> So I just tried the latest notebook environment in Interactive mode. It seems to be a bit slow in initializing but it works. No extra libraries needed…I just import the tensorflow library.</p>\n<p>In the first run of the notebook it starts to load (I guess? what else…) my roughly 200GB of TFRecords from GCS. Initially it crashed on out of memory and the kernel restart…however after that it runs just fine.</p>\n<p>When the model starts training it does however give the following message multiple times ( seems to be on each Callback for on_epoch_end).<br>\nOverall interactive works…I now Queued it. Will let you know once it has finished if it went OK.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F0661fbc29cbd90cdc1ab843eccea865b%2Ferror.png?generation=1680724135065055&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi @herbison So I just tried the latest notebook environment in Interactive mode. It seems to be a bit slow in initializing but it works. No extra libraries needed...I just import the tensorflow library.\n\nIn the first run of the notebook it starts to load (I guess? what else...) my roughly 200GB of TFRecords from GCS. Initially it crashed on out of memory and the kernel restart...however after that it runs just fine.\n\nWhen the model starts training it does however give the following message multiple times ( seems to be on each Callback for on_epoch_end).\nOverall interactive works...I now Queued it. Will let you know once it has finished if it went OK.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F0661fbc29cbd90cdc1ab843eccea865b%2Ferror.png?generation=1680724135065055&alt=media)",
      "votes": null
    },
    {
      "id": "2211112",
      "postDate": "04/05/2023 19:55:56",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Glad to know its working! Thanks for the update :)</p>\n<p>btw, you don't need to load from GCS (though you can) for TPU VMs.<br>\nYou can read directly from the filesystem (ex. if you have a Kaggle dataset in /kaggle/input). Though since you have such a large dataset, it may not be available on kaggle already and GCS might be the right way to work with it in that case.</p>",
      "rawMarkdown": "rsmits Glad to know its working! Thanks for the update :)\n\nbtw, you don't need to load from GCS (though you can) for TPU VMs.\nYou can read directly from the filesystem (ex. if you have a Kaggle dataset in /kaggle/input). Though since you have such a large dataset, it may not be available on kaggle already and GCS might be the right way to work with it in that case.",
      "votes": null
    },
    {
      "id": "2211163",
      "postDate": "04/05/2023 20:55:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Yes as far as dataset is concerned I have to use GCS because of the size. This way it is also easier to use in combination with Colab.</p>\n<p>I just tried the Save and Run option for the notebook. Again a crash on the first attempt. On the second attempt I lowered the amount of files to load to around 175GB total. That way it started running.</p>\n<p>Still a bit surprised by this…with the 24 february 2023 notebook version and TF 2.9.1 I was able to run the same code with more than 200GB loaded.<br>\nNot sure if this is something that others experience though.</p>\n<p>Anyway in this stage of the competition I'am happy that it works and I have ways to work around this issue.</p>\n<p>Thanks for your time …for me it solved and working again :-)</p>",
      "rawMarkdown": "Hi @herbison Yes as far as dataset is concerned I have to use GCS because of the size. This way it is also easier to use in combination with Colab.\n\nI just tried the Save and Run option for the notebook. Again a crash on the first attempt. On the second attempt I lowered the amount of files to load to around 175GB total. That way it started running.\n\nStill a bit surprised by this...with the 24 february 2023 notebook version and TF 2.9.1 I was able to run the same code with more than 200GB loaded.\nNot sure if this is something that others experience though.\n\nAnyway in this stage of the competition I'am happy that it works and I have ways to work around this issue.\n\nThanks for your time ...for me it solved and working again :-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2208039,
      "author_name": "herbison",
      "author_url": "",
      "post_date": "04/03/2023 18:51:45",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> The \"latest\" tf (v130 March 29 2023) was rolled back because of bug with juptyer-lsp locking up the entire notebook for TPU VMs. This was the first build of the TPU VM that came with tf for TPU VM pre-installed (vs. needing to be installed).</p>\n<p>I have a fixed version (v131 Apr 3 2023) which is unreleased but seems to fix the lsp issue and otherwise be identical to the v130 release. Hoping for that to be released early this week.</p>\n<p>In the meantime, you'll need to pip install the tf tpu vm package to get TPU VM speedup.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2208053,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/03/2023 19:01:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Thanks for the quick reply. Ok I will keep the above in mind.</p>\n<p>With the pip install for the tf tpu vm package you mean uncomment the line that installs the TF 2.9.1 version?</p>\n<p>Or keep the default TF 2.11 and install what specific package?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2208061,
              "author_name": "herbison",
              "author_url": "",
              "post_date": "04/03/2023 19:06:26",
              "content": "<p>It would be uncomment to install tf 2.9.1, you can also install other tpu vm versions of tensorflow, but you'll need to update the libtpu version as well, example for tf 2.11:<br>\n<a href=\"https://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm\" target=\"_blank\">https://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm</a></p>\n<p>There's a table of versions:<br>\n<a href=\"https://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm\" target=\"_blank\">https://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2208090,
                  "author_name": "rsmits",
                  "author_url": "",
                  "post_date": "04/03/2023 19:43:10",
                  "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Ok tried both of the above…with TF 2.9.1 and the lines for installing TF 2.11. With some delays .. the notebook gets to the point where it would start training. Then show different kinds of 'model_pruner' errors and then seems to hang or reboot the kernel.</p>\n<p>I think I'll wait for the updated notebook version to be released as this is killing my TPU quota. Thanks for the support anyway.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2208092,
                      "author_name": "herbison",
                      "author_url": "",
                      "post_date": "04/03/2023 19:44:31",
                      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Feel free to share the notebook with me too and I can evaluate it.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2208116,
                          "author_name": "rsmits",
                          "author_url": "",
                          "post_date": "04/03/2023 19:57:50",
                          "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> ok send you a direct message with some questions regarding that.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2210668,
                              "author_name": "herbison",
                              "author_url": "",
                              "post_date": "04/05/2023 14:46:28",
                              "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> The \"latest\" TPU VM image should have tf2.12 for TPU VM preinstalled correctly.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2210733,
                                  "author_name": "rsmits",
                                  "author_url": "",
                                  "post_date": "04/05/2023 15:32:30",
                                  "content": "<p><a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Ok thanks for giving a notice. I will try later today and let you know the results !</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2211100,
                                      "author_name": "rsmits",
                                      "author_url": "",
                                      "post_date": "04/05/2023 19:48:58",
                                      "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> So I just tried the latest notebook environment in Interactive mode. It seems to be a bit slow in initializing but it works. No extra libraries needed…I just import the tensorflow library.</p>\n<p>In the first run of the notebook it starts to load (I guess? what else…) my roughly 200GB of TFRecords from GCS. Initially it crashed on out of memory and the kernel restart…however after that it runs just fine.</p>\n<p>When the model starts training it does however give the following message multiple times ( seems to be on each Callback for on_epoch_end).<br>\nOverall interactive works…I now Queued it. Will let you know once it has finished if it went OK.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F0661fbc29cbd90cdc1ab843eccea865b%2Ferror.png?generation=1680724135065055&amp;alt=media\" alt=\"\"></p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2211112,
                                          "author_name": "herbison",
                                          "author_url": "",
                                          "post_date": "04/05/2023 19:55:56",
                                          "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Glad to know its working! Thanks for the update :)</p>\n<p>btw, you don't need to load from GCS (though you can) for TPU VMs.<br>\nYou can read directly from the filesystem (ex. if you have a Kaggle dataset in /kaggle/input). Though since you have such a large dataset, it may not be available on kaggle already and GCS might be the right way to work with it in that case.</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 2211163,
                                              "author_name": "rsmits",
                                              "author_url": "",
                                              "post_date": "04/05/2023 20:55:06",
                                              "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> Yes as far as dataset is concerned I have to use GCS because of the size. This way it is also easier to use in combination with Colab.</p>\n<p>I just tried the Save and Run option for the notebook. Again a crash on the first attempt. On the second attempt I lowered the amount of files to load to around 175GB total. That way it started running.</p>\n<p>Still a bit surprised by this…with the 24 february 2023 notebook version and TF 2.9.1 I was able to run the same code with more than 200GB loaded.<br>\nNot sure if this is something that others experience though.</p>\n<p>Anyway in this stage of the competition I'am happy that it works and I have ways to work around this issue.</p>\n<p>Thanks for your time …for me it solved and working again :-)</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2207991": "Not sure if anybody else is experiencing issues with the TPU VM accelerators? My training notebook that used to work on Saturday now seems to be not using the TPU.....the speed achieved looks more like a very descent desktop ;-)\n\nI get the errors below. I tried with the default TF 2.11 installed and also with the install of TF 2.9.1...both give the same errors.\n\nWith only 20 hours of TPU time quota there is not a lot of room for troubleshooting.\n\n@herbison Mentioning you because I'am not sure if Kaggle knows this error?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2Fee602c4fd6b90bd35362b80aedb5ae7d%2Ftpu.png?generation=1680546335585738&alt=media)\n\nUpdate: The same error happens on the Example notebook that is provided on the Product Feedback page:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F2e92b7ba14ca0dea57eefbe9833b9ec1%2Ftpu2.png?generation=1680546901060758&alt=media)",
    "2208039": "rsmits The \"latest\" tf (v130 March 29 2023) was rolled back because of bug with juptyer-lsp locking up the entire notebook for TPU VMs. This was the first build of the TPU VM that came with tf for TPU VM pre-installed (vs. needing to be installed).\n\nI have a fixed version (v131 Apr 3 2023) which is unreleased but seems to fix the lsp issue and otherwise be identical to the v130 release. Hoping for that to be released early this week.\n\nIn the meantime, you'll need to pip install the tf tpu vm package to get TPU VM speedup.",
    "2208053": "Hi @herbison Thanks for the quick reply. Ok I will keep the above in mind.\n\nWith the pip install for the tf tpu vm package you mean uncomment the line that installs the TF 2.9.1 version?\n\nOr keep the default TF 2.11 and install what specific package?",
    "2208061": "It would be uncomment to install tf 2.9.1, you can also install other tpu vm versions of tensorflow, but you'll need to update the libtpu version as well, example for tf 2.11:\nhttps://www.kaggle.com/code/herbison/tensorflow-2-11-on-tpu-vm\n\nThere's a table of versions:\nhttps://cloud.google.com/tpu/docs/supported-tpu-configurations#tpu_vm",
    "2208090": "herbison Ok tried both of the above...with TF 2.9.1 and the lines for installing TF 2.11. With some delays .. the notebook gets to the point where it would start training. Then show different kinds of 'model_pruner' errors and then seems to hang or reboot the kernel.\n\nI think I'll wait for the updated notebook version to be released as this is killing my TPU quota. Thanks for the support anyway.",
    "2208092": "rsmits Feel free to share the notebook with me too and I can evaluate it.",
    "2208116": "herbison ok send you a direct message with some questions regarding that.",
    "2210668": "rsmits The \"latest\" TPU VM image should have tf2.12 for TPU VM preinstalled correctly.",
    "2210733": "herbison Ok thanks for giving a notice. I will try later today and let you know the results !",
    "2211100": "Hi @herbison So I just tried the latest notebook environment in Interactive mode. It seems to be a bit slow in initializing but it works. No extra libraries needed...I just import the tensorflow library.\n\nIn the first run of the notebook it starts to load (I guess? what else...) my roughly 200GB of TFRecords from GCS. Initially it crashed on out of memory and the kernel restart...however after that it runs just fine.\n\nWhen the model starts training it does however give the following message multiple times ( seems to be on each Callback for on_epoch_end).\nOverall interactive works...I now Queued it. Will let you know once it has finished if it went OK.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1073620%2F0661fbc29cbd90cdc1ab843eccea865b%2Ferror.png?generation=1680724135065055&alt=media)",
    "2211112": "rsmits Glad to know its working! Thanks for the update :)\n\nbtw, you don't need to load from GCS (though you can) for TPU VMs.\nYou can read directly from the filesystem (ex. if you have a Kaggle dataset in /kaggle/input). Though since you have such a large dataset, it may not be available on kaggle already and GCS might be the right way to work with it in that case.",
    "2211163": "Hi @herbison Yes as far as dataset is concerned I have to use GCS because of the size. This way it is also easier to use in combination with Colab.\n\nI just tried the Save and Run option for the notebook. Again a crash on the first attempt. On the second attempt I lowered the amount of files to load to around 175GB total. That way it started running.\n\nStill a bit surprised by this...with the 24 february 2023 notebook version and TF 2.9.1 I was able to run the same code with more than 200GB loaded.\nNot sure if this is something that others experience though.\n\nAnyway in this stage of the competition I'am happy that it works and I have ways to work around this issue.\n\nThanks for your time ...for me it solved and working again :-)"
  },
  "source": "meta"
}