{
  "id": 514484,
  "title": "TPU Error?",
  "url": "/competitions/leash-BELKA/discussion/514484",
  "author_name": "",
  "post_date": "2024-06-24T08:55:26.520736900Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>When I runing my model on TPU, machine, it gets an error in the function  <code>fit()</code> , which seems to indicate that a device does not exist, but my TPU has already been initialized normally. How can I solve this problem? Thanks!!</p>\n<pre><code>ValueError                                Traceback (most recent  last)\nCell In[], line \n      model = NewDCNN_model()\n      model.summary()\n---&gt;  history = model.fit(\n          X_train,\n          y_train,\n          validation_data=(X_val, y_val),\n          epochs=CFG.EPOCHS,\n          callbacks=[checkpoint, reduce_lr_loss, estimate],\n          batch_size=CFG.BATCH_SIZE,\n          verbose=\n      )\n      model.load_weights(f)\n      oof = model.predict(X_val, batch_size = *CFG.BATCH_SIZE)\n\n localpython3.kerastils/traceback_utils.py:, in filter_traceback.&lt;locals&gt;.error_handler(*args, **kwargs)\n         filtered_tb = _process_traceback_frames(e.__traceback__)\n         # To get the full stack trace, :\n         # `keras.config.disable_traceback_filtering()`\n--&gt;      raise e.with_traceback(filtered_tb)  None\n     :\n         del filtered_tb\n\n localpython3.tensorflowdistribute/packed_distributed_variable.py:, in PackedDistributedVariable.get_var_on_device(self, device)\n         d == device:\n           self._distributed_variables[i]\n---&gt;  raise ValueError( % device)\n\nValueError: Device replica:device:CPU: is not found\n</code></pre>",
  "messages": [
    {
      "id": "2887538",
      "postDate": "06/24/2024 08:55:26",
      "content": "<p>When I runing my model on TPU, machine, it gets an error in the function  <code>fit()</code> , which seems to indicate that a device does not exist, but my TPU has already been initialized normally. How can I solve this problem? Thanks!!</p>\n<pre><code>ValueError                                Traceback (most recent  last)\nCell In[], line \n      model = NewDCNN_model()\n      model.summary()\n---&gt;  history = model.fit(\n          X_train,\n          y_train,\n          validation_data=(X_val, y_val),\n          epochs=CFG.EPOCHS,\n          callbacks=[checkpoint, reduce_lr_loss, estimate],\n          batch_size=CFG.BATCH_SIZE,\n          verbose=\n      )\n      model.load_weights(f)\n      oof = model.predict(X_val, batch_size = *CFG.BATCH_SIZE)\n\n localpython3.kerastils/traceback_utils.py:, in filter_traceback.&lt;locals&gt;.error_handler(*args, **kwargs)\n         filtered_tb = _process_traceback_frames(e.__traceback__)\n         # To get the full stack trace, :\n         # `keras.config.disable_traceback_filtering()`\n--&gt;      raise e.with_traceback(filtered_tb)  None\n     :\n         del filtered_tb\n\n localpython3.tensorflowdistribute/packed_distributed_variable.py:, in PackedDistributedVariable.get_var_on_device(self, device)\n         d == device:\n           self._distributed_variables[i]\n---&gt;  raise ValueError( % device)\n\nValueError: Device replica:device:CPU: is not found\n</code></pre>",
      "rawMarkdown": "When I runing my model on TPU, machine, it gets an error in the function  `fit()` , which seems to indicate that a device does not exist, but my TPU has already been initialized normally. How can I solve this problem? Thanks!!\n\n```\nValueError                                Traceback (most recent call last)\nCell In[12], line 47\n     44 model = NewDCNN_model()\n     45 model.summary()\n---> 47 history = model.fit(\n     48     X_train,\n     49     y_train,\n     50     validation_data=(X_val, y_val),\n     51     epochs=CFG.EPOCHS,\n     52     callbacks=[checkpoint, reduce_lr_loss, estimate],\n     53     batch_size=CFG.BATCH_SIZE,\n     54     verbose=1\n     55 )\n     56 model.load_weights(f\"model-{fold}.weights.h5\")\n     57 oof = model.predict(X_val, batch_size = 2*CFG.BATCH_SIZE)\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:122, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    119     filtered_tb = _process_traceback_frames(e.__traceback__)\n    120     # To get the full stack trace, call:\n    121     # `keras.config.disable_traceback_filtering()`\n--> 122     raise e.with_traceback(filtered_tb) from None\n    123 finally:\n    124     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/tensorflow/python/distribute/packed_distributed_variable.py:91, in PackedDistributedVariable.get_var_on_device(self, device)\n     89   if d == device:\n     90     return self._distributed_variables[i]\n---> 91 raise ValueError(\"Device %s is not found\" % device)\n\nValueError: Device /job:localhost/replica:0/task:0/device:CPU:0 is not found\n```",
      "votes": null
    },
    {
      "id": "2892109",
      "postDate": "06/27/2024 05:15:46",
      "content": "<p>having the same issue, one epoch runs just fine after which I get errors. I am now running multiple epochs like this, not ideal so do let me know if you find the proper solution </p>\n<p>for _ in range(5):<br>\n        try:<br>\n            history = model.fit(<br>\n                    X_train, y_train,<br>\n                    validation_data=(X_val, y_val),<br>\n                    epochs=1,<br>\n                    callbacks=[reduce_lr_loss],<br>\n                    batch_size=CFG.BATCH_SIZE,<br>\n                    verbose=1,<br>\n                )<br>\n        except:<br>\n            print()           </p>",
      "rawMarkdown": "having the same issue, one epoch runs just fine after which I get errors. I am now running multiple epochs like this, not ideal so do let me know if you find the proper solution \n     \nfor _ in range(5):\n        try:\n            history = model.fit(\n                    X_train, y_train,\n                    validation_data=(X_val, y_val),\n                    epochs=1,\n                    callbacks=[reduce_lr_loss],\n                    batch_size=CFG.BATCH_SIZE,\n                    verbose=1,\n                )\n        except:\n            print()",
      "votes": null
    },
    {
      "id": "2892453",
      "postDate": "06/27/2024 08:01:17",
      "content": "<p>Yes, I am also facing the same problem, but so far I have not been able to find an effective solution.😞</p>",
      "rawMarkdown": "Yes, I am also facing the same problem, but so far I have not been able to find an effective solution.😞",
      "votes": null
    },
    {
      "id": "2898535",
      "postDate": "07/01/2024 07:52:50",
      "content": "<p>Have you tried this?  : </p>\n<pre><code>:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=) \n    \n    \n    \n    strategy = tf.distribute.TPUStrategy(tpu)\n    ()\n    (, strategy.num_replicas_in_sync)\n:\n    strategy = tf.distribute.get_strategy()\n</code></pre>",
      "rawMarkdown": "Have you tried this?  : \n```python\n\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    #tf.config.experimental_connect_to_cluster(tpu)\n    #tf.tpu.experimental.initialize_tpu_system(tpu)\n    #.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"Running on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()\n ```",
      "votes": null
    },
    {
      "id": "2900425",
      "postDate": "07/02/2024 09:16:32",
      "content": "<p>Easy solution: copy <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a>, and load your notebook into it, so you'll be able to use old environment (not latest). For me, 2 gpus works only in latest, while tpu in this old environment.<br>\nMost likely something with tensorflow or keras versions.</p>",
      "rawMarkdown": "Easy solution: copy [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data), and load your notebook into it, so you'll be able to use old environment (not latest). For me, 2 gpus works only in latest, while tpu in this old environment.\nMost likely something with tensorflow or keras versions.",
      "votes": null
    },
    {
      "id": "2900444",
      "postDate": "07/02/2024 09:40:15",
      "content": "<p>Okay, I will try. Thx! 🙏🏼</p>",
      "rawMarkdown": "Okay, I will try. Thx! 🙏🏼",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2892109,
      "author_name": "mayankdeshwal",
      "author_url": "",
      "post_date": "06/27/2024 05:15:46",
      "content": "<p>having the same issue, one epoch runs just fine after which I get errors. I am now running multiple epochs like this, not ideal so do let me know if you find the proper solution </p>\n<p>for _ in range(5):<br>\n        try:<br>\n            history = model.fit(<br>\n                    X_train, y_train,<br>\n                    validation_data=(X_val, y_val),<br>\n                    epochs=1,<br>\n                    callbacks=[reduce_lr_loss],<br>\n                    batch_size=CFG.BATCH_SIZE,<br>\n                    verbose=1,<br>\n                )<br>\n        except:<br>\n            print()           </p>",
      "votes": null,
      "replies": [
        {
          "id": 2892453,
          "author_name": "lau01b",
          "author_url": "",
          "post_date": "06/27/2024 08:01:17",
          "content": "<p>Yes, I am also facing the same problem, but so far I have not been able to find an effective solution.😞</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2898535,
          "author_name": "lau01b",
          "author_url": "",
          "post_date": "07/01/2024 07:52:50",
          "content": "<p>Have you tried this?  : </p>\n<pre><code>:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=) \n    \n    \n    \n    strategy = tf.distribute.TPUStrategy(tpu)\n    ()\n    (, strategy.num_replicas_in_sync)\n:\n    strategy = tf.distribute.get_strategy()\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2900425,
      "author_name": "yermvad",
      "author_url": "",
      "post_date": "07/02/2024 09:16:32",
      "content": "<p>Easy solution: copy <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a>, and load your notebook into it, so you'll be able to use old environment (not latest). For me, 2 gpus works only in latest, while tpu in this old environment.<br>\nMost likely something with tensorflow or keras versions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2900444,
          "author_name": "lau01b",
          "author_url": "",
          "post_date": "07/02/2024 09:40:15",
          "content": "<p>Okay, I will try. Thx! 🙏🏼</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2887538": "When I runing my model on TPU, machine, it gets an error in the function  `fit()` , which seems to indicate that a device does not exist, but my TPU has already been initialized normally. How can I solve this problem? Thanks!!\n\n```\nValueError                                Traceback (most recent call last)\nCell In[12], line 47\n     44 model = NewDCNN_model()\n     45 model.summary()\n---> 47 history = model.fit(\n     48     X_train,\n     49     y_train,\n     50     validation_data=(X_val, y_val),\n     51     epochs=CFG.EPOCHS,\n     52     callbacks=[checkpoint, reduce_lr_loss, estimate],\n     53     batch_size=CFG.BATCH_SIZE,\n     54     verbose=1\n     55 )\n     56 model.load_weights(f\"model-{fold}.weights.h5\")\n     57 oof = model.predict(X_val, batch_size = 2*CFG.BATCH_SIZE)\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:122, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    119     filtered_tb = _process_traceback_frames(e.__traceback__)\n    120     # To get the full stack trace, call:\n    121     # `keras.config.disable_traceback_filtering()`\n--> 122     raise e.with_traceback(filtered_tb) from None\n    123 finally:\n    124     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/tensorflow/python/distribute/packed_distributed_variable.py:91, in PackedDistributedVariable.get_var_on_device(self, device)\n     89   if d == device:\n     90     return self._distributed_variables[i]\n---> 91 raise ValueError(\"Device %s is not found\" % device)\n\nValueError: Device /job:localhost/replica:0/task:0/device:CPU:0 is not found\n```",
    "2892109": "having the same issue, one epoch runs just fine after which I get errors. I am now running multiple epochs like this, not ideal so do let me know if you find the proper solution \n     \nfor _ in range(5):\n        try:\n            history = model.fit(\n                    X_train, y_train,\n                    validation_data=(X_val, y_val),\n                    epochs=1,\n                    callbacks=[reduce_lr_loss],\n                    batch_size=CFG.BATCH_SIZE,\n                    verbose=1,\n                )\n        except:\n            print()",
    "2892453": "Yes, I am also facing the same problem, but so far I have not been able to find an effective solution.😞",
    "2898535": "Have you tried this?  : \n```python\n\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    #tf.config.experimental_connect_to_cluster(tpu)\n    #tf.tpu.experimental.initialize_tpu_system(tpu)\n    #.connect(tpu=\"local\") # \"local\" for 1VM TPU\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(\"Running on TPU\")\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\nexcept:\n    strategy = tf.distribute.get_strategy()\n ```",
    "2900425": "Easy solution: copy [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data), and load your notebook into it, so you'll be able to use old environment (not latest). For me, 2 gpus works only in latest, while tpu in this old environment.\nMost likely something with tensorflow or keras versions.",
    "2900444": "Okay, I will try. Thx! 🙏🏼"
  },
  "source": "meta"
}