{
  "id": 557233,
  "title": "Early Error busts code?",
  "url": "/competitions/tpu-getting-started/discussion/557233",
  "author_name": "Monica Backes",
  "post_date": "2025-01-17T16:59:38.118000",
  "votes": 0,
  "comment_count": 2,
  "views": null,
  "content": "<p>Hi everyone, I haven't made any changes to the \"Create your first submission\" code, but after executing the first cell, which only imports a few libraries and prints the TF version, I get the following error:</p>\n<p>WARNING: Logging before InitGoogle() is written to STDERR<br>\nE0000 00:00:1737131892.857967      13 common_lib.cc:818] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in <code>tpu_process_addresses</code>=\"local\"<br>\n=== Source Location Trace: ===<br>\nlearning/45eac/tfrc/runtime/common_lib.cc:501<br>\nTensorflow version 2.15.0</p>\n<p>After the next cell I get output telling me that the TPU is running with 8 cores so I assume all to be happy.</p>\n<p>However, when it comes to fit the model, I get the following error, which seems to connect to the first, and I can go no further.</p>\n<p>WARNING: All log messages before absl::InitializeLog() is called are written to STDERR<br>\nI0000 00:00:1737132157.191984      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.<br>\n2025-01-17 16:42:37.196554: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.237753: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.276614: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.314997: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.353776: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.392098: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.431183: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</p>\n<h2>2025-01-17 16:42:37.469038: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</h2>\n<p>RuntimeError                              Traceback (most recent call last)<br>\nCell In[17], line 5<br>\n      2 EPOCHS = 12<br>\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n----&gt; 5 history = model.fit(<br>\n      6     ds_train,<br>\n      7     validation_data=ds_valid,<br>\n      8     epochs=EPOCHS,<br>\n      9     steps_per_epoch=STEPS_PER_EPOCH,<br>\n     10     callbacks=[lr_callback],<br>\n     11 )</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback..error_handler(*args, **kwargs)<br>\n    120     filtered_tb = _process_traceback_frames(e.<strong>traceback</strong>)<br>\n    121     # To get the full stack trace, call:<br>\n    122     # <code>keras.config.disable_traceback_filtering()</code><br>\n--&gt; 123     raise e.with_traceback(filtered_tb) from None<br>\n    124 finally:<br>\n    125     del filtered_tb</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)<br>\n     27 else:<br>\n     28     colocate_var = reference_variable<br>\n---&gt; 30 with self._distribution_strategy.extended.colocate_vars_with(<br>\n     31     colocate_var<br>\n     32 ):<br>\n     33     return super().add_variable_from_reference(<br>\n     34         reference_variable, name=name, initializer=initializer<br>\n     35     )</p>\n<p>RuntimeError: Mixing different tf.distribute.Strategy objects:  is not </p>\n<p>Is anyone able to explain what is going on here, and how I correct the error?  This is my first foray into TPUs and CNNs.</p>\n<p>Thanks!</p>\n<p>Monica</p>",
  "messages": [
    {
      "id": 3197806,
      "postDate": "2025-05-08T15:57:29.553Z",
      "content": "<p>Hey! I have the same error. Did you find the solution yet?</p>",
      "rawMarkdown": "Hey! I have the same error. Did you find the solution yet?",
      "replies": [
        {
          "id": 3201703,
          "postDate": "2025-05-14T09:33:19.573Z",
          "content": "<p>I'm afraid not.  I spent some time happily working just with GPUs, and have only recently come back to this.  </p>",
          "rawMarkdown": "I'm afraid not.  I spent some time happily working just with GPUs, and have only recently come back to this.  "
        }
      ]
    },
    {
      "id": 3099382,
      "postDate": "2025-01-17T16:59:38.120Z",
      "content": "<p>Hi everyone, I haven't made any changes to the \"Create your first submission\" code, but after executing the first cell, which only imports a few libraries and prints the TF version, I get the following error:</p>\n<p>WARNING: Logging before InitGoogle() is written to STDERR<br>\nE0000 00:00:1737131892.857967      13 common_lib.cc:818] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in <code>tpu_process_addresses</code>=\"local\"<br>\n=== Source Location Trace: ===<br>\nlearning/45eac/tfrc/runtime/common_lib.cc:501<br>\nTensorflow version 2.15.0</p>\n<p>After the next cell I get output telling me that the TPU is running with 8 cores so I assume all to be happy.</p>\n<p>However, when it comes to fit the model, I get the following error, which seems to connect to the first, and I can go no further.</p>\n<p>WARNING: All log messages before absl::InitializeLog() is called are written to STDERR<br>\nI0000 00:00:1737132157.191984      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.<br>\n2025-01-17 16:42:37.196554: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.237753: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.276614: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.314997: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.353776: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.392098: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2025-01-17 16:42:37.431183: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</p>\n<h2>2025-01-17 16:42:37.469038: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</h2>\n<p>RuntimeError                              Traceback (most recent call last)<br>\nCell In[17], line 5<br>\n      2 EPOCHS = 12<br>\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n----&gt; 5 history = model.fit(<br>\n      6     ds_train,<br>\n      7     validation_data=ds_valid,<br>\n      8     epochs=EPOCHS,<br>\n      9     steps_per_epoch=STEPS_PER_EPOCH,<br>\n     10     callbacks=[lr_callback],<br>\n     11 )</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback..error_handler(*args, **kwargs)<br>\n    120     filtered_tb = _process_traceback_frames(e.<strong>traceback</strong>)<br>\n    121     # To get the full stack trace, call:<br>\n    122     # <code>keras.config.disable_traceback_filtering()</code><br>\n--&gt; 123     raise e.with_traceback(filtered_tb) from None<br>\n    124 finally:<br>\n    125     del filtered_tb</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)<br>\n     27 else:<br>\n     28     colocate_var = reference_variable<br>\n---&gt; 30 with self._distribution_strategy.extended.colocate_vars_with(<br>\n     31     colocate_var<br>\n     32 ):<br>\n     33     return super().add_variable_from_reference(<br>\n     34         reference_variable, name=name, initializer=initializer<br>\n     35     )</p>\n<p>RuntimeError: Mixing different tf.distribute.Strategy objects:  is not </p>\n<p>Is anyone able to explain what is going on here, and how I correct the error?  This is my first foray into TPUs and CNNs.</p>\n<p>Thanks!</p>\n<p>Monica</p>",
      "rawMarkdown": "Hi everyone, I haven't made any changes to the \"Create your first submission\" code, but after executing the first cell, which only imports a few libraries and prints the TF version, I get the following error:\n\nWARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1737131892.857967      13 common_lib.cc:818] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:501\nTensorflow version 2.15.0\n\nAfter the next cell I get output telling me that the TPU is running with 8 cores so I assume all to be happy.\n\nHowever, when it comes to fit the model, I get the following error, which seems to connect to the first, and I can go no further.\n\nWARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nI0000 00:00:1737132157.191984      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.\n2025-01-17 16:42:37.196554: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.237753: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.276614: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.314997: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.353776: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.392098: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.431183: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.469038: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\nCell In[17], line 5\n      2 EPOCHS = 12\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n----> 5 history = model.fit(\n      6     ds_train,\n      7     validation_data=ds_valid,\n      8     epochs=EPOCHS,\n      9     steps_per_epoch=STEPS_PER_EPOCH,\n     10     callbacks=[lr_callback],\n     11 )\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    120     filtered_tb = _process_traceback_frames(e.__traceback__)\n    121     # To get the full stack trace, call:\n    122     # `keras.config.disable_traceback_filtering()`\n--> 123     raise e.with_traceback(filtered_tb) from None\n    124 finally:\n    125     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)\n     27 else:\n     28     colocate_var = reference_variable\n---> 30 with self._distribution_strategy.extended.colocate_vars_with(\n     31     colocate_var\n     32 ):\n     33     return super().add_variable_from_reference(\n     34         reference_variable, name=name, initializer=initializer\n     35     )\n\nRuntimeError: Mixing different tf.distribute.Strategy objects: <tensorflow.python.distribute.tpu_strategy.TPUStrategyV2 object at 0x7d2ce1f9ae00> is not <tensorflow.python.distribute.distribute_lib._DefaultDistributionStrategy object at 0x7d24142043d0>\n\nIs anyone able to explain what is going on here, and how I correct the error?  This is my first foray into TPUs and CNNs.\n\nThanks!\n\nMonica"
    }
  ],
  "comments": [
    {
      "id": 3197806,
      "author_name": "Mark Slavin",
      "author_url": "",
      "post_date": "2025-05-08T15:57:29.553000",
      "content": "<p>Hey! I have the same error. Did you find the solution yet?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3201703,
          "author_name": "Monica Backes",
          "author_url": "",
          "post_date": "2025-05-14T09:33:19.573000",
          "content": "<p>I'm afraid not.  I spent some time happily working just with GPUs, and have only recently come back to this.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3197806": "Hey! I have the same error. Did you find the solution yet?",
    "3099382": "Hi everyone, I haven't made any changes to the \"Create your first submission\" code, but after executing the first cell, which only imports a few libraries and prints the TF version, I get the following error:\n\nWARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1737131892.857967      13 common_lib.cc:818] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:501\nTensorflow version 2.15.0\n\nAfter the next cell I get output telling me that the TPU is running with 8 cores so I assume all to be happy.\n\nHowever, when it comes to fit the model, I get the following error, which seems to connect to the first, and I can go no further.\n\nWARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nI0000 00:00:1737132157.191984      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.\n2025-01-17 16:42:37.196554: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.237753: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.276614: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.314997: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.353776: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.392098: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.431183: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2025-01-17 16:42:37.469038: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\nCell In[17], line 5\n      2 EPOCHS = 12\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n----> 5 history = model.fit(\n      6     ds_train,\n      7     validation_data=ds_valid,\n      8     epochs=EPOCHS,\n      9     steps_per_epoch=STEPS_PER_EPOCH,\n     10     callbacks=[lr_callback],\n     11 )\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    120     filtered_tb = _process_traceback_frames(e.__traceback__)\n    121     # To get the full stack trace, call:\n    122     # `keras.config.disable_traceback_filtering()`\n--> 123     raise e.with_traceback(filtered_tb) from None\n    124 finally:\n    125     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)\n     27 else:\n     28     colocate_var = reference_variable\n---> 30 with self._distribution_strategy.extended.colocate_vars_with(\n     31     colocate_var\n     32 ):\n     33     return super().add_variable_from_reference(\n     34         reference_variable, name=name, initializer=initializer\n     35     )\n\nRuntimeError: Mixing different tf.distribute.Strategy objects: <tensorflow.python.distribute.tpu_strategy.TPUStrategyV2 object at 0x7d2ce1f9ae00> is not <tensorflow.python.distribute.distribute_lib._DefaultDistributionStrategy object at 0x7d24142043d0>\n\nIs anyone able to explain what is going on here, and how I correct the error?  This is my first foray into TPUs and CNNs.\n\nThanks!\n\nMonica"
  }
}