{
  "id": 476956,
  "title": "error during model training",
  "url": "/competitions/tpu-getting-started/discussion/476956",
  "author_name": "Serhii Kotliar",
  "post_date": "2024-02-14T06:52:27.760000",
  "votes": 6,
  "comment_count": 6,
  "views": null,
  "content": "<p>Hello colleagues. This error occurs when training the model. Please advise what to do?</p>\n<h1>Define training epochs</h1>\n<p>EPOCHS = 12<br>\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n​<br>\nhistory = model.fit(<br>\n    ds_train,<br>\n    validation_data=ds_valid,<br>\n    epochs=EPOCHS,<br>\n    steps_per_epoch=STEPS_PER_EPOCH,<br>\n    callbacks=[lr_callback],<br>\n)</p>\n<p>WARNING: All log messages before absl::InitializeLog() is called are written to STDERR<br>\nI0000 00:00:1707739625.233410      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.<br>\n2024-02-12 12:07:05.236619: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.276507: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.318318: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.359488: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.398781: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.436990: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.475228: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</p>\n<h2>2024-02-12 12:07:05.513166: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</h2>\n<p>RuntimeError                              Traceback (most recent call last)<br>\nCell In[16], line 5<br>\n      2 EPOCHS = 12<br>\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n----&gt; 5 history = model.fit(<br>\n      6     ds_train,<br>\n      7     validation_data=ds_valid,<br>\n      8     epochs=EPOCHS,<br>\n      9     steps_per_epoch=STEPS_PER_EPOCH,<br>\n     10     callbacks=[lr_callback],<br>\n     11 )</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback..error_handler(*args, **kwargs)<br>\n    120     filtered_tb = _process_traceback_frames(e.<strong>traceback</strong>)<br>\n    121     # To get the full stack trace, call:<br>\n    122     # <code>keras.config.disable_traceback_filtering()</code><br>\n--&gt; 123     raise e.with_traceback(filtered_tb) from None<br>\n    124 finally:<br>\n    125     del filtered_tb</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)<br>\n     27 else:<br>\n     28     colocate_var = reference_variable<br>\n---&gt; 30 with self._distribution_strategy.extended.colocate_vars_with(<br>\n     31     colocate_var<br>\n     32 ):<br>\n     33     return super().add_variable_from_reference(<br>\n     34         reference_variable, name=name, initializer=initializer<br>\n     35     )</p>\n<p>RuntimeError: Mixing different tf.distribute.Strategy objects:  is not </p>\n<h4>in Google Colab with TPU, Tensorflow version 2.12 is connected and there is no error, but in Kaggle Tensorflow version 2.15 is connected and an error appears. This means that the code in the notepad is already outdated and needs to be updated.</h4>",
  "messages": [
    {
      "id": 2651426,
      "postDate": "2024-02-14T06:52:27.760Z",
      "content": "<p>Hello colleagues. This error occurs when training the model. Please advise what to do?</p>\n<h1>Define training epochs</h1>\n<p>EPOCHS = 12<br>\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n​<br>\nhistory = model.fit(<br>\n    ds_train,<br>\n    validation_data=ds_valid,<br>\n    epochs=EPOCHS,<br>\n    steps_per_epoch=STEPS_PER_EPOCH,<br>\n    callbacks=[lr_callback],<br>\n)</p>\n<p>WARNING: All log messages before absl::InitializeLog() is called are written to STDERR<br>\nI0000 00:00:1707739625.233410      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.<br>\n2024-02-12 12:07:05.236619: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.276507: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.318318: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.359488: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.398781: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.436990: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.<br>\n2024-02-12 12:07:05.475228: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</p>\n<h2>2024-02-12 12:07:05.513166: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.</h2>\n<p>RuntimeError                              Traceback (most recent call last)<br>\nCell In[16], line 5<br>\n      2 EPOCHS = 12<br>\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE<br>\n----&gt; 5 history = model.fit(<br>\n      6     ds_train,<br>\n      7     validation_data=ds_valid,<br>\n      8     epochs=EPOCHS,<br>\n      9     steps_per_epoch=STEPS_PER_EPOCH,<br>\n     10     callbacks=[lr_callback],<br>\n     11 )</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback..error_handler(*args, **kwargs)<br>\n    120     filtered_tb = _process_traceback_frames(e.<strong>traceback</strong>)<br>\n    121     # To get the full stack trace, call:<br>\n    122     # <code>keras.config.disable_traceback_filtering()</code><br>\n--&gt; 123     raise e.with_traceback(filtered_tb) from None<br>\n    124 finally:<br>\n    125     del filtered_tb</p>\n<p>File /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)<br>\n     27 else:<br>\n     28     colocate_var = reference_variable<br>\n---&gt; 30 with self._distribution_strategy.extended.colocate_vars_with(<br>\n     31     colocate_var<br>\n     32 ):<br>\n     33     return super().add_variable_from_reference(<br>\n     34         reference_variable, name=name, initializer=initializer<br>\n     35     )</p>\n<p>RuntimeError: Mixing different tf.distribute.Strategy objects:  is not </p>\n<h4>in Google Colab with TPU, Tensorflow version 2.12 is connected and there is no error, but in Kaggle Tensorflow version 2.15 is connected and an error appears. This means that the code in the notepad is already outdated and needs to be updated.</h4>",
      "rawMarkdown": "Hello colleagues. This error occurs when training the model. Please advise what to do?\n\n# Define training epochs\nEPOCHS = 12\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n​\nhistory = model.fit(\n    ds_train,\n    validation_data=ds_valid,\n    epochs=EPOCHS,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    callbacks=[lr_callback],\n)\n\n\nWARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nI0000 00:00:1707739625.233410      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.\n2024-02-12 12:07:05.236619: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.276507: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.318318: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.359488: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.398781: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.436990: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.475228: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.513166: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\nCell In[16], line 5\n      2 EPOCHS = 12\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n----> 5 history = model.fit(\n      6     ds_train,\n      7     validation_data=ds_valid,\n      8     epochs=EPOCHS,\n      9     steps_per_epoch=STEPS_PER_EPOCH,\n     10     callbacks=[lr_callback],\n     11 )\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    120     filtered_tb = _process_traceback_frames(e.__traceback__)\n    121     # To get the full stack trace, call:\n    122     # `keras.config.disable_traceback_filtering()`\n--> 123     raise e.with_traceback(filtered_tb) from None\n    124 finally:\n    125     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)\n     27 else:\n     28     colocate_var = reference_variable\n---> 30 with self._distribution_strategy.extended.colocate_vars_with(\n     31     colocate_var\n     32 ):\n     33     return super().add_variable_from_reference(\n     34         reference_variable, name=name, initializer=initializer\n     35     )\n\nRuntimeError: Mixing different tf.distribute.Strategy objects: <tensorflow.python.distribute.tpu_strategy.TPUStrategy object at 0x7cb5ab4fb9a0> is not <tensorflow.python.distribute.distribute_lib._DefaultDistributionStrategy object at 0x7cac506e35b0>\n\n####  in Google Colab with TPU, Tensorflow version 2.12 is connected and there is no error, but in Kaggle Tensorflow version 2.15 is connected and an error appears. This means that the code in the notepad is already outdated and needs to be updated.",
      "votes": 6
    },
    {
      "id": 2655397,
      "postDate": "2024-02-16T20:35:05.930Z",
      "content": "<p>The support service does not want to answer this question. He says that the group should answer, but no one answers 😁</p>",
      "rawMarkdown": "The support service does not want to answer this question. He says that the group should answer, but no one answers 😁",
      "votes": 1
    },
    {
      "id": 2654976,
      "postDate": "2024-02-16T15:56:59.257Z",
      "content": "<p>Got the same error, interested too</p>",
      "rawMarkdown": "Got the same error, interested too",
      "votes": 1
    },
    {
      "id": 2661961,
      "postDate": "2024-02-21T16:06:14.547Z",
      "content": "<p>I think this is happening becauase TPU is not getting allocated. using GPU works. <br>\ntry: # detect TPUs<br>\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection<br>\n    strategy = tf.distribute.TPUStrategy(tpu)<br>\nexcept ValueError: # detect GPUs<br>\n    strategy = tf.distribute.MirroredStrategy() # for GPU or multi-GPU machines<br>\n    #strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU<br>\n    #strategy = tf.distribute.experimental.MultiWorkerMirroredStrategy() # for clusters of multi-GPU machines</p>\n<p>print(\"Number of accelerators: \", strategy.num_replicas_in_sync)</p>\n<p>I took the above from: <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n\" target=\"_blank\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n</a></p>",
      "rawMarkdown": "I think this is happening becauase TPU is not getting allocated. using GPU works. \ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.MirroredStrategy() # for GPU or multi-GPU machines\n    #strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\n    #strategy = tf.distribute.experimental.MultiWorkerMirroredStrategy() # for clusters of multi-GPU machines\n\nprint(\"Number of accelerators: \", strategy.num_replicas_in_sync)\n\nI took the above from: https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n",
      "replies": [
        {
          "id": 2662081,
          "postDate": "2024-02-21T17:33:51.057Z",
          "content": "<p>Unfortunately, it doesn’t work in Kaggle, but it really works in Google Colab. But the reason is not clear.</p>",
          "rawMarkdown": "Unfortunately, it doesn’t work in Kaggle, but it really works in Google Colab. But the reason is not clear.",
          "replies": [
            {
              "id": 2663168,
              "postDate": "2024-02-22T10:55:31.960Z",
              "content": "<p>I was able to run on Kaggle. but you need to reduce image size from 512x512 to 224x224 and also reduce the batch size to 4 or 8. This is because lower memory availability </p>",
              "rawMarkdown": "I was able to run on Kaggle. but you need to reduce image size from 512x512 to 224x224 and also reduce the batch size to 4 or 8. This is because lower memory availability "
            },
            {
              "id": 2663541,
              "postDate": "2024-02-22T14:53:20.767Z",
              "content": "<p>I checked all these options, but still this error seems to be in Kaggle🤔</p>",
              "rawMarkdown": "I checked all these options, but still this error seems to be in Kaggle🤔"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2655397,
      "author_name": "Serhii Kotliar",
      "author_url": "",
      "post_date": "2024-02-16T20:35:05.930000",
      "content": "<p>The support service does not want to answer this question. He says that the group should answer, but no one answers 😁</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2654976,
      "author_name": "LcRslt",
      "author_url": "",
      "post_date": "2024-02-16T15:56:59.257000",
      "content": "<p>Got the same error, interested too</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2661961,
      "author_name": "Haridas",
      "author_url": "",
      "post_date": "2024-02-21T16:06:14.547000",
      "content": "<p>I think this is happening becauase TPU is not getting allocated. using GPU works. <br>\ntry: # detect TPUs<br>\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection<br>\n    strategy = tf.distribute.TPUStrategy(tpu)<br>\nexcept ValueError: # detect GPUs<br>\n    strategy = tf.distribute.MirroredStrategy() # for GPU or multi-GPU machines<br>\n    #strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU<br>\n    #strategy = tf.distribute.experimental.MultiWorkerMirroredStrategy() # for clusters of multi-GPU machines</p>\n<p>print(\"Number of accelerators: \", strategy.num_replicas_in_sync)</p>\n<p>I took the above from: <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n\" target=\"_blank\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2662081,
          "author_name": "Serhii Kotliar",
          "author_url": "",
          "post_date": "2024-02-21T17:33:51.057000",
          "content": "<p>Unfortunately, it doesn’t work in Kaggle, but it really works in Google Colab. But the reason is not clear.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2663168,
              "author_name": "Haridas",
              "author_url": "",
              "post_date": "2024-02-22T10:55:31.960000",
              "content": "<p>I was able to run on Kaggle. but you need to reduce image size from 512x512 to 224x224 and also reduce the batch size to 4 or 8. This is because lower memory availability </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663541,
              "author_name": "Serhii Kotliar",
              "author_url": "",
              "post_date": "2024-02-22T14:53:20.767000",
              "content": "<p>I checked all these options, but still this error seems to be in Kaggle🤔</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2651426": "Hello colleagues. This error occurs when training the model. Please advise what to do?\n\n# Define training epochs\nEPOCHS = 12\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n​\nhistory = model.fit(\n    ds_train,\n    validation_data=ds_valid,\n    epochs=EPOCHS,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    callbacks=[lr_callback],\n)\n\n\nWARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nI0000 00:00:1707739625.233410      13 device_compiler.h:186] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.\n2024-02-12 12:07:05.236619: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.276507: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.318318: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.359488: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.398781: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.436990: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.475228: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n2024-02-12 12:07:05.513166: E external/local_xla/xla/stream_executor/stream_executor_internal.h:177] SetPriority unimplemented for this stream.\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\nCell In[16], line 5\n      2 EPOCHS = 12\n      3 STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n----> 5 history = model.fit(\n      6     ds_train,\n      7     validation_data=ds_valid,\n      8     epochs=EPOCHS,\n      9     steps_per_epoch=STEPS_PER_EPOCH,\n     10     callbacks=[lr_callback],\n     11 )\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs)\n    120     filtered_tb = _process_traceback_frames(e.__traceback__)\n    121     # To get the full stack trace, call:\n    122     # `keras.config.disable_traceback_filtering()`\n--> 123     raise e.with_traceback(filtered_tb) from None\n    124 finally:\n    125     del filtered_tb\n\nFile /usr/local/lib/python3.10/site-packages/keras/src/backend/tensorflow/optimizer.py:30, in TFOptimizer.add_variable_from_reference(self, reference_variable, name, initializer)\n     27 else:\n     28     colocate_var = reference_variable\n---> 30 with self._distribution_strategy.extended.colocate_vars_with(\n     31     colocate_var\n     32 ):\n     33     return super().add_variable_from_reference(\n     34         reference_variable, name=name, initializer=initializer\n     35     )\n\nRuntimeError: Mixing different tf.distribute.Strategy objects: <tensorflow.python.distribute.tpu_strategy.TPUStrategy object at 0x7cb5ab4fb9a0> is not <tensorflow.python.distribute.distribute_lib._DefaultDistributionStrategy object at 0x7cac506e35b0>\n\n####  in Google Colab with TPU, Tensorflow version 2.12 is connected and there is no error, but in Kaggle Tensorflow version 2.15 is connected and an error appears. This means that the code in the notepad is already outdated and needs to be updated.",
    "2655397": "The support service does not want to answer this question. He says that the group should answer, but no one answers 😁",
    "2654976": "Got the same error, interested too",
    "2661961": "I think this is happening becauase TPU is not getting allocated. using GPU works. \ntry: # detect TPUs\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect() # TPU detection\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError: # detect GPUs\n    strategy = tf.distribute.MirroredStrategy() # for GPU or multi-GPU machines\n    #strategy = tf.distribute.get_strategy() # default strategy that works on CPU and single GPU\n    #strategy = tf.distribute.experimental.MultiWorkerMirroredStrategy() # for clusters of multi-GPU machines\n\nprint(\"Number of accelerators: \", strategy.num_replicas_in_sync)\n\nI took the above from: https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/07_Keras_Flowers_TPU_xception_fine_tuned_best.ipynb#scrollTo=FpvUOuC3j27n"
  }
}