{
  "id": 278161,
  "title": "TPU Enabled Model Runs on CPU",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/278161",
  "author_name": "",
  "post_date": "2021-10-13T03:23:51.361102700Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I’m trying to convert an interesting notebook that I found to run on the TPU. </p>\n<p>However, when I run model.fit() the CPU meter spikes, while the TPU stays at zero, and the epoch wants to run for over 20 minutes. </p>\n<p>How can I troubleshoot the problem?</p>\n<p><a href=\"https://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu\" target=\"_blank\">https://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu</a></p>",
  "messages": [
    {
      "id": "1542977",
      "postDate": "10/13/2021 03:23:51",
      "content": "<p>I’m trying to convert an interesting notebook that I found to run on the TPU. </p>\n<p>However, when I run model.fit() the CPU meter spikes, while the TPU stays at zero, and the epoch wants to run for over 20 minutes. </p>\n<p>How can I troubleshoot the problem?</p>\n<p><a href=\"https://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu\" target=\"_blank\">https://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu</a></p>",
      "rawMarkdown": "I’m trying to convert an interesting notebook that I found to run on the TPU. \n\nHowever, when I run model.fit() the CPU meter spikes, while the TPU stays at zero, and the epoch wants to run for over 20 minutes. \n\nHow can I troubleshoot the problem?\n\nhttps://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu",
      "votes": null
    },
    {
      "id": "1543305",
      "postDate": "10/13/2021 10:34:51",
      "content": "<p>Well, I've just opened this in edit mode and the Accelerator is set to <code>GPU</code> right now. <br>\nAlso I see that you are reading the files from the disk as you go, this will end up being a CPU bound process especially at first. The TPU will work but only very briefly as it quickly processes the batch of data read in by the relatively slow CPU.</p>",
      "rawMarkdown": "Well, I've just opened this in edit mode and the Accelerator is set to `GPU` right now. \nAlso I see that you are reading the files from the disk as you go, this will end up being a CPU bound process especially at first. The TPU will work but only very briefly as it quickly processes the batch of data read in by the relatively slow CPU.",
      "votes": null
    },
    {
      "id": "1543333",
      "postDate": "10/13/2021 11:26:01",
      "content": "<p>I probably was working on the GPU when I saved and only switched DEVICE = \"TPU\" before the save.  I did another Quick Save to fix.  Thanks.</p>\n<p>Yes, there's a lot of data.  Was hoping to switch to the png's to speed up the load.</p>\n<p>Anyway, this value needs to be set and the proper Accelerator Setting to match.</p>\n<pre><code>DEVICE = \"TPU\" # \"TPU\" #or \"GPU\"\n</code></pre>\n<p>In GPU mode, it runs much faster:</p>\n<pre><code>=== Training flair ===\nDownloading data from https://storage.googleapis.com/keras-applications/efficientnetb3_notop.h5\n43941888/43941136 [==============================] - 0s 0us/step\n#########################\nTraining...\nEpoch 1/5\n15/15 - 129s - loss: 0.7304 - binary_accuracy: 0.5863\nEpoch 2/5\n</code></pre>\n<p>In TPU mode after the Epoch begins I am told it will take about 20 minutes per, (when I don't get an unavailable error)</p>\n<pre><code>#########################\n=== Training flair ===\n#########################\nTraining...\nEpoch 1/5\n[Some message about 20 minutes]\n</code></pre>\n<p>By the way, this morning I am getting this \"Unavailable error\" quite a bit.  TPU's are at capacity?</p>\n<pre><code>\":[{\"created\":\"@1634124060.293603685\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n     [[{{node MultiDeviceIteratorGetNextFromShard}}]]\n     [[RemoteCall]]\n     [[IteratorGetNextAsOptional]]\n     [[Cast_1/_62]]\n  (4) Unavailable: {{function_node __inference_train_function_158782}} f ... [truncated]\n</code></pre>\n<p>Also, the original author used a \"generator\", which I'm not familiar with.  Could be a problem?</p>\n<pre><code>    history = model.fit(\n        train_generator,\n        epochs=epochs,\n        steps_per_epoch=len(train_generator),\n        verbose=2,\n        workers=2\n    )    \n</code></pre>\n<p>The old API looked like this:</p>\n<pre><code>    history = model.fit_generator(\n        generator=train_generator,\n        steps_per_epoch=len(train_generator),\n        epochs=epochs,\n        workers=2\n    )\n</code></pre>",
      "rawMarkdown": "I probably was working on the GPU when I saved and only switched DEVICE = \"TPU\" before the save.  I did another Quick Save to fix.  Thanks.\n\nYes, there's a lot of data.  Was hoping to switch to the png's to speed up the load.\n\nAnyway, this value needs to be set and the proper Accelerator Setting to match.\n\n```\nDEVICE = \"TPU\" # \"TPU\" #or \"GPU\"\n```\n\nIn GPU mode, it runs much faster:\n\n```\n=== Training flair ===\nDownloading data from https://storage.googleapis.com/keras-applications/efficientnetb3_notop.h5\n43941888/43941136 [==============================] - 0s 0us/step\n#########################\nTraining...\nEpoch 1/5\n15/15 - 129s - loss: 0.7304 - binary_accuracy: 0.5863\nEpoch 2/5\n```\n\nIn TPU mode after the Epoch begins I am told it will take about 20 minutes per, (when I don't get an unavailable error)\n\n```\n#########################\n=== Training flair ===\n#########################\nTraining...\nEpoch 1/5\n[Some message about 20 minutes]\n```\n\nBy the way, this morning I am getting this \"Unavailable error\" quite a bit.  TPU's are at capacity?\n\n```\n\":[{\"created\":\"@1634124060.293603685\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n\t [[{{node MultiDeviceIteratorGetNextFromShard}}]]\n\t [[RemoteCall]]\n\t [[IteratorGetNextAsOptional]]\n\t [[Cast_1/_62]]\n  (4) Unavailable: {{function_node __inference_train_function_158782}} f ... [truncated]\n```\n\nAlso, the original author used a \"generator\", which I'm not familiar with.  Could be a problem?\n\n```\n    history = model.fit(\n        train_generator,\n        epochs=epochs,\n        steps_per_epoch=len(train_generator),\n        verbose=2,\n        workers=2\n    )    \n```\n\nThe old API looked like this:\n\n```\n    history = model.fit_generator(\n        generator=train_generator,\n        steps_per_epoch=len(train_generator),\n        epochs=epochs,\n        workers=2\n    )\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1543305,
      "author_name": "mutantspore",
      "author_url": "",
      "post_date": "10/13/2021 10:34:51",
      "content": "<p>Well, I've just opened this in edit mode and the Accelerator is set to <code>GPU</code> right now. <br>\nAlso I see that you are reading the files from the disk as you go, this will end up being a CPU bound process especially at first. The TPU will work but only very briefly as it quickly processes the batch of data read in by the relatively slow CPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1543333,
          "author_name": "mmellinger66",
          "author_url": "",
          "post_date": "10/13/2021 11:26:01",
          "content": "<p>I probably was working on the GPU when I saved and only switched DEVICE = \"TPU\" before the save.  I did another Quick Save to fix.  Thanks.</p>\n<p>Yes, there's a lot of data.  Was hoping to switch to the png's to speed up the load.</p>\n<p>Anyway, this value needs to be set and the proper Accelerator Setting to match.</p>\n<pre><code>DEVICE = \"TPU\" # \"TPU\" #or \"GPU\"\n</code></pre>\n<p>In GPU mode, it runs much faster:</p>\n<pre><code>=== Training flair ===\nDownloading data from https://storage.googleapis.com/keras-applications/efficientnetb3_notop.h5\n43941888/43941136 [==============================] - 0s 0us/step\n#########################\nTraining...\nEpoch 1/5\n15/15 - 129s - loss: 0.7304 - binary_accuracy: 0.5863\nEpoch 2/5\n</code></pre>\n<p>In TPU mode after the Epoch begins I am told it will take about 20 minutes per, (when I don't get an unavailable error)</p>\n<pre><code>#########################\n=== Training flair ===\n#########################\nTraining...\nEpoch 1/5\n[Some message about 20 minutes]\n</code></pre>\n<p>By the way, this morning I am getting this \"Unavailable error\" quite a bit.  TPU's are at capacity?</p>\n<pre><code>\":[{\"created\":\"@1634124060.293603685\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n     [[{{node MultiDeviceIteratorGetNextFromShard}}]]\n     [[RemoteCall]]\n     [[IteratorGetNextAsOptional]]\n     [[Cast_1/_62]]\n  (4) Unavailable: {{function_node __inference_train_function_158782}} f ... [truncated]\n</code></pre>\n<p>Also, the original author used a \"generator\", which I'm not familiar with.  Could be a problem?</p>\n<pre><code>    history = model.fit(\n        train_generator,\n        epochs=epochs,\n        steps_per_epoch=len(train_generator),\n        verbose=2,\n        workers=2\n    )    \n</code></pre>\n<p>The old API looked like this:</p>\n<pre><code>    history = model.fit_generator(\n        generator=train_generator,\n        steps_per_epoch=len(train_generator),\n        epochs=epochs,\n        workers=2\n    )\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1542977": "I’m trying to convert an interesting notebook that I found to run on the TPU. \n\nHowever, when I run model.fit() the CPU meter spikes, while the TPU stays at zero, and the epoch wants to run for over 20 minutes. \n\nHow can I troubleshoot the problem?\n\nhttps://www.kaggle.com/mmellinger66/brain-tumor-keras-efficientnet-tpu",
    "1543305": "Well, I've just opened this in edit mode and the Accelerator is set to `GPU` right now. \nAlso I see that you are reading the files from the disk as you go, this will end up being a CPU bound process especially at first. The TPU will work but only very briefly as it quickly processes the batch of data read in by the relatively slow CPU.",
    "1543333": "I probably was working on the GPU when I saved and only switched DEVICE = \"TPU\" before the save.  I did another Quick Save to fix.  Thanks.\n\nYes, there's a lot of data.  Was hoping to switch to the png's to speed up the load.\n\nAnyway, this value needs to be set and the proper Accelerator Setting to match.\n\n```\nDEVICE = \"TPU\" # \"TPU\" #or \"GPU\"\n```\n\nIn GPU mode, it runs much faster:\n\n```\n=== Training flair ===\nDownloading data from https://storage.googleapis.com/keras-applications/efficientnetb3_notop.h5\n43941888/43941136 [==============================] - 0s 0us/step\n#########################\nTraining...\nEpoch 1/5\n15/15 - 129s - loss: 0.7304 - binary_accuracy: 0.5863\nEpoch 2/5\n```\n\nIn TPU mode after the Epoch begins I am told it will take about 20 minutes per, (when I don't get an unavailable error)\n\n```\n#########################\n=== Training flair ===\n#########################\nTraining...\nEpoch 1/5\n[Some message about 20 minutes]\n```\n\nBy the way, this morning I am getting this \"Unavailable error\" quite a bit.  TPU's are at capacity?\n\n```\n\":[{\"created\":\"@1634124060.293603685\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n\t [[{{node MultiDeviceIteratorGetNextFromShard}}]]\n\t [[RemoteCall]]\n\t [[IteratorGetNextAsOptional]]\n\t [[Cast_1/_62]]\n  (4) Unavailable: {{function_node __inference_train_function_158782}} f ... [truncated]\n```\n\nAlso, the original author used a \"generator\", which I'm not familiar with.  Could be a problem?\n\n```\n    history = model.fit(\n        train_generator,\n        epochs=epochs,\n        steps_per_epoch=len(train_generator),\n        verbose=2,\n        workers=2\n    )    \n```\n\nThe old API looked like this:\n\n```\n    history = model.fit_generator(\n        generator=train_generator,\n        steps_per_epoch=len(train_generator),\n        epochs=epochs,\n        workers=2\n    )\n```"
  },
  "source": "meta"
}